9 days ago
London, United KingdomSenior
Responsibilities
- Build and evolve platform infrastructure as code using Terraform, Kubernetes, and GCP.
- Own CI/CD pipelines end-to-end to make builds, tests, and deployments fast, reliable, and secure.
- Own observability systems covering metrics, logging, and distributed tracing.
- Define and maintain SLOs for latency, availability, and error rates, and manage capacity against them.
- Lead incident response, resolve complex production issues, run blameless post-mortems, and implement fixes as code.
- Partner with product engineering teams to reduce manual toil and improve the path from code creation to production operation.
Requirements
- Hands-on experience operating production systems with an SRE mindset, including SLOs, error budgets, toil reduction, and blameless post-mortems.
- Strong knowledge of modern cloud architectures using GCP or AWS, including networking, IAM, and managed services.
- Deep expertise with Docker and Kubernetes at scale, including cluster design, networking, storage, and security.
- Proficiency in a modern language such as TypeScript, Go, Python, or Rust, with experience writing well-tested automation.
- Experience owning mission-critical operational tooling, including centralized logging or APM tools such as Datadog or Splunk.
- Ability to communicate technical ideas and influence how engineers build and operate services.
Benefits
- Competitive salary of £105,000 to £125,000.
- Equity in an early-stage technology company.
- 25 days of holiday plus local public holidays.
- Apple hardware.
- Private medical insurance through AXA.
- Pension contribution through Hargreaves Lansdown.
- Enhanced family leave.
- Team off-sites in various locations.
