9 days ago
London, United KingdomSenior
Responsibilities
- Build and evolve reproducible, version-controlled platform infrastructure using Terraform, Kubernetes, and GCP.
- Own CI/CD pipelines end-to-end to make builds, tests, and deployments fast, reliable, and secure.
- Own metrics, logging, and distributed tracing through the observability stack.
- Define and maintain SLOs for latency, availability, and error rates while managing capacity.
- Lead production incident response, resolve complex issues, run blameless post-mortems, and implement fixes as code.
- Partner with product engineering teams to reduce manual toil and improve the path from code creation to production operation.
Requirements
- Hands-on experience operating production systems with an SRE mindset, including SLOs, error budgets, toil reduction, and blameless post-mortems.
- Strong knowledge of modern cloud architectures on GCP or AWS, including networking, IAM, and managed services.
- Deep expertise operating Docker and Kubernetes at scale, including cluster design, networking, storage, and security.
- Proficiency in a modern language such as Typescript, Go, Python, or Rust, with experience writing well-tested automation.
- Experience owning mission-critical operational tooling, including centralized logging or APM such as Datadog or Splunk.
- Ability to communicate technical ideas and influence how other engineers build and operate their services.
Benefits
- Competitive salary of £105,000 to £125,000.
- Equity in an early-stage technology company.
- 25 days of holiday plus local public holidays.
- Apple hardware.
- Private medical insurance through AXA.
- Pension contribution through Hargreaves Lansdown.
- Enhanced family leave.
- Team off-sites in locations including Barcelona, Lisbon, Malta, and Split.
