1 month ago
San Diego, CA, USAMid Level
Base Salary
$125k - $145k/yr
Responsibilities
- Build, maintain, and optimize CI/CD pipelines for application deployments.
- Manage and improve AWS infrastructure across ECS, EKS, and EC2 environments.
- Develop and maintain Infrastructure as Code using Terraform.
- Monitor system health and performance using Grafana and supporting observability tools.
- Troubleshoot infrastructure and application issues in staging and production.
- Collaborate with developers to improve deployment workflows and system reliability.
- Support containerized workloads using Docker and Kubernetes.
- Implement and maintain logging, alerting, and incident response processes.
- Contribute to AWS cost optimization and performance tuning.
- Maintain infrastructure documentation, processes, and runbooks.
- Use AI-assisted tooling and emerging technologies to improve productivity, infrastructure automation, CI/CD optimization, incident investigation, observability, and operational efficiency.
- Independently handle infrastructure and deployment tasks, take ownership of services, and identify reliability, performance, and cost improvements.
Requirements
- 2–4 years of experience in DevOps, SRE, or related roles.
- Hands-on experience with AWS, particularly ECS, EKS, and EC2.
- Experience writing and maintaining Terraform-based Infrastructure as Code.
- Experience with monitoring and observability tools, especially Grafana.
- Familiarity with GitHub Actions, GitLab CI, Jenkins, or similar CI/CD tools.
- Experience with Docker and containerization.
- Working knowledge of Linux systems and networking fundamentals.
- Scripting experience with Bash, Python, or similar languages.
- Ability to debug across application and infrastructure layers and escalate complex architectural decisions appropriately.
- Experience applying AI to DevOps practices, including infrastructure automation, CI/CD optimization, incident investigation, observability, and operational efficiency.
- Preferred experience includes operating production Kubernetes clusters, using Prometheus or other Grafana metrics backends, AWS IAM and networking, load balancing, autoscaling, high availability, fault-tolerant design, blue/green or canary deployments, and AWS cost monitoring and optimization.
Benefits
- Comprehensive health, dental, vision, retirement matching, paid time off, parental leave, and an annual education and training stipend for full-time permanent employees.
- Hybrid work model with three days per week in the San Diego office.
- Weekly lunches, quality coffee, social events, and location-dependent amenities such as parent rooms, gyms, lounges, and adaptable workstations.
- Wellness workshops, events, and a complimentary Headspace subscription.
- Employee resource groups and an inclusive, collaborative workplace.
- Full-time, permanent employment with a multi-stage interview process.
