13 hours ago
Seattle, WA, USA +2 moreSenior
Base Salary
$140k - $274k/yr
Responsibilities
- Design scalable, fault-tolerant infrastructure across AWS, GCP, and Azure.
- Operate production container platforms using Kubernetes, Helm, and infrastructure-as-code tooling.
- Automate operational tasks and infrastructure management with Python or Go.
- Own reliability, performance, and efficiency for core services, including SLOs, error budgets, and on-call operations.
- Lead incident response, post-mortems, root-cause analyses, and preventative architectural improvements.
- Shape short- and long-term platform, observability, cost, and reliability initiatives.
- Collaborate with product, security, and engineering teams on reliable and scalable system design.
- Use and expand AI-assisted workflows for incident investigation, infrastructure changes, runbooks, tooling, and code review.
Requirements
- At least 5 years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, high-availability production systems.
- Production experience running containerized workloads in real clusters, with Helm and Terraform or Pulumi on at least one major cloud provider.
- Strong proficiency in Python or Go for automation and tooling.
- Existing daily experience with AI-assisted workflows and tools such as Claude Code, Droid, Codex, or comparable internal tooling.
- Ability to identify systemic reliability weaknesses, evaluate tradeoffs, and design innovative solutions.
- Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent tools.
- Excellent communication, collaboration, problem-solving, and cross-functional relationship-building skills.
- Demonstrated ownership of at least one end-to-end 0-to-1 infrastructure build with measurable outcomes.
- Preferred: software-engineering experience building and shipping non-trivial production services, libraries, or internal frameworks in Python, Go, or a comparable language.
Benefits
- Hybrid work from the New York City, San Francisco, Seattle, or London hubs.
- Generous paid time off and company holidays.
- Medical, dental, and vision coverage for employees and families.
- Sixteen weeks of paid parental leave for all parents.
- Fertility and family planning support.
- Early-detection cancer testing through Galleri.
- Flexible spending accounts, dependent FSA options, and eligible-plan HSA contributions.
- Wellness, learning and development, and work-life stipends.
- Company-wide and team off-sites.
- Company stock options and a 401(k).
Tech Stack
Categories
DevOpsSite Reliability
