9 months ago
Base Salary
$230k - $310k/yr
Responsibilities
- Own the reliability, availability, and performance of production systems across AWS infrastructure.
- Build metrics, logging, tracing, and alerting infrastructure to improve system observability.
- Design and ship automation that reduces toil, makes deployments safer, and accelerates recovery from incidents.
- Lead incident response and blameless post-mortems and implement systemic fixes.
- Partner with engineering teams on architecture reviews, SLO and SLI design, and scalable reliability practices.
- Manage and optimize compute, networking, databases, and managed services.
Requirements
- 5+ years of experience in site reliability engineering, DevOps, or systems engineering.
- Deep, hands-on AWS expertise and strong programming skills in Python, Go, or TypeScript/Node.js.
- Experience with Terraform or CloudFormation and end-to-end observability solutions.
- Track record of improving system reliability through automation, monitoring, and architectural improvements.
- Deep understanding of networking, distributed systems, Docker, Kubernetes, and database performance at scale.
- Strong incident management and debugging skills for complex production failures.
- Nice-to-have experience scaling SaaS products to millions of users or working with Kafka, chaos engineering, or service mesh technologies.
- Nice-to-have AWS certifications or experience with SOC 2 or ISO 27001 security and compliance frameworks.
Benefits
- Full-time position with benefits and equity.
- Strong in-office culture in San Francisco, working in person 4–5 days per week, with flexibility to work from home when focus matters most.
Categories
DevOpsSite Reliability
About Gamma
Gamma is your AI design partner for creating presentations, websites, social media, and more. All at the speed of thought.
