About Arena Club
Arena Club is building the number one digital marketplace for collectibles. We started by reinventing the hobby with the first-ever digital card show, and we are expanding into new categories so collectors can buy, sell, grade, and showcase everything they care about in one place. Our ‘Slab Pack’ product brought the mystery and chase of opening physical packs of cards to a digital format, and has led to Arena Club being one of the fastest-growing (and highly profitable) companies in Los Angeles. Our business was founded by 5x World Series Champion Derek Jeter and serial entrepreneur Brian Lee, who previously founded LegalZoom and The Honest Company, both of which are now public companies.
About The Role
We’re seeking a Senior DevOps Engineer to design, build, and operate the infrastructure, deployment systems, and operational tooling that support Arena Club’s platform. In this role, you will own critical components of our cloud and delivery environment, ensuring systems are secure, reliable, scalable, and efficient. You will drive infrastructure and automation initiatives end-to-end, improving system performance, deployment velocity, and developer productivity.
This position is designed for engineers who operate independently, consistently deliver high-quality solutions, and proactively identify opportunities to improve reliability, scalability, and operational efficiency. As a trusted partner to engineering and product teams, you will help enable fast, safe, and resilient product delivery in a fast-paced environment where your work directly impacts system uptime, performance, and overall platform health.
What You Will Do
- Design, implement, and maintain CI/CD pipelines for web, backend, and mobile applications to support reliable and repeatable deployment
- Manage and optimize Kubernetes clusters to ensure high availability, performance, and scalability in a production environment
- Build and maintain infrastructure as code using Terraform across the AWS environment
- Design and implement a secure, scalable, and cost-effective cloud architecture leveraging AWS services
- Develop automation scripts and internal tooling to reduce manual work and improve operational efficiency
- Establish and maintain advanced monitoring, logging, alerting, and incident response practices using DataDog, CloudWatch, and Cloudflare
- Troubleshoot production incidents, perform root cause analysis, and implement long-term reliability improvements
- Identify infrastructure risks, performance bottlenecks, and technical debt, and drive proactive remediation
- Partner closely with development teams to improve deployment workflows, system observability, and operational readiness
- Promote DevOps best practices, reliability standards, and operational ownership across the Engineering team
- Document infrastructure architecture, operational procedures, and system dependencies to support team continuity and resilience
Qualifications
- 6+ years of experience in DevOps, Site Reliability Engineering, platform engineering, or infrastructure-focused software engineering roles
- Strong experience designing and operating cloud infrastructure in AWS production environments
- Hands-on experience managing Kubernetes clusters at scale
- Proven experience building and maintaining CI/CD pipelines
- Strong experience with infrastructure as code tools such as Terraform
- Experience implementing monitoring, logging, and observability solutions in production systems
- Proficiency in scripting or programming (Python, Bash, or similar)
- Strong understanding of cloud networking, security best practices, and system reliability principles
- Ability to independently own infrastructure initiatives from design through implementation and operation
- Strong problem-solving skills and the ability to communicate technical risks, trade-offs, and operational impacts effectively
- Ability to operate effectively in a fast-paced startup environment while maintaining system stability and high operational standards
Preferred Qualifications
- Experience supporting high-growth or high-availability production environments
- Experience with cost optimization and performance tuning in AWS
- Experience driving incident management, on-call practices, or reliability improvements
- Experience mentoring engineers or influencing DevOps culture across teams
Tech Stack
- Cloud & Infrastructure: AWS (EC2, EKS, S3, RDS, ECR, etc.)
- Containerization & Orchestration: Docker, Kubernetes
- Package Management & Artifact Repository: Helm, ChartMuseum, ECR
- Infrastructure as Code: Terraform
- CI/CD: GitHub Actions, CircleCI or similar pipeline tooling
- Secrets & Configuration Management: AWS SSm, AWS Secrets Manager, External Secret Operator
- Monitoring & Security: DataDog, Sentry, Cloudflare, AWS CloudWatch
- Scripting & Automation: Python, Bash