13 hours ago
London, United KingdomSenior
Responsibilities
- Design scalable, fault-tolerant infrastructure across AWS, GCP, and Azure.
- Automate infrastructure management and operational tasks using Python or Go.
- Operate Kubernetes, Helm, Terraform, cloud tooling, and internal platform surfaces.
- Own reliability, performance, efficiency, SLOs, error budgets, and on-call operations for core services.
- Lead incident response, post-mortems, root-cause analyses, and preventative architecture improvements.
- Shape multi-year platform investments in observability, cost management, and reliability.
- Collaborate with product, security, and engineering teams on reliable system design from conception through launch.
- Use and develop AI-assisted workflows for incident investigation, infrastructure changes, runbooks, tooling, and code review.
Requirements
- At least 5 years of experience in infrastructure engineering, DevOps, or a similar role operating large-scale, high-availability production systems.
- Production experience running containerized workloads and real clusters.
- Experience with Helm and Terraform or Pulumi on at least one major cloud, preferably AWS.
- Strong proficiency in Python or Go for automation and tooling.
- Existing daily experience with AI-assisted engineering workflows, including agentic tooling adoption.
- Experience with monitoring and logging stacks such as Prometheus, Grafana, and ELK or equivalent.
- Demonstrated ability to identify systemic reliability weaknesses, make tradeoffs, and solve complex infrastructure problems.
- Excellent communication, collaboration, ownership, and problem-solving skills.
- Experience owning at least one infrastructure build end-to-end with measurable outcomes.
- Preferred: software engineering experience building and shipping non-trivial production services, libraries, or internal frameworks in Python, Go, or a comparable language.
Benefits
- Hybrid work based out of the London or New York City hubs.
- Generous paid time off and company holidays.
- Comprehensive medical and dental insurance.
- 16 weeks of paid parental leave for all parents.
- Fertility and family planning support.
- Early-detection cancer testing through Galleri.
- Competitive pension scheme and company contribution.
- Wellness, learning and development, and annual work-life stipends.
- Company-wide and team off-sites.
- Company stock options and competitive compensation.
Tech Stack
Categories
DevOpsSite Reliability
