
Site Reliability Engineer
Capital on Tap2 months ago
London, United KingdomSenior
Responsibilities
- Manage and automate Azure, Datadog, NGINX, and Cloudflare.
- Develop and monitor Kubernetes and serverless resources.
- Maintain infrastructure code using Terraform, CRDs, and Crossplane.
- Improve platform performance, systems, processes, and stakeholder-facing solutions.
- Contribute to application architecture and design processes.
- Automate repetitive tasks and reduce operational toil.
- Create SLIs and SLOs and increase application visibility.
- Align with the Product team on SLAs and core service objectives.
- Collaborate with Platform Engineers on automated solutions and pipelines.
- Optimize infrastructure and pipelines to improve user experience.
- Support Azure DevOps, Octopus Deploy, and Flux for software delivery.
- Lead incident troubleshooting to protect customer experience.
Requirements
- Experience managing a public cloud, with Azure experience advantageous.
- Experience with Azure DevOps, Octopus, Flux, or other CI/CD tools.
- Experience with Linux and Microsoft systems.
- Strong communication and collaboration skills across multiple teams.
- Proficiency with infrastructure-as-code technologies, including Terraform or Pulumi.
- Experience with cloud monitoring solutions, with Datadog experience advantageous.
- Experience with Kubernetes and Docker.
- Experience with at least one scripting language: Python, PowerShell, or Go.
Benefits
- Private healthcare including dental and optician services through Vitality.
- Worldwide travel insurance through Vitality.
- Anniversary rewards of £250, £500, £750, and a four-week fully paid sabbatical.
- Salary sacrifice pension scheme with up to 7% matching.
- 28 days of holiday plus bank holidays.
- Annual learning and wellbeing budget.
- Enhanced parental leave.
- Cycle to Work Scheme.
- Season Ticket Loan.
- Six free therapy sessions per year.
- Dog-friendly offices with free drinks and snacks, a pool table, arcade machine, beer tap, and office dogs.
- Hybrid work arrangement based in London with two days in the office.
Categories
Site Reliability