13 hours ago
London, United KingdomSenior
Responsibilities
- Eliminate operational toil through automation, software development, and internal SaaS tools.
- Partner with application teams and internal stakeholders to improve reliability and standardization.
- Build and scale a resilient, cost-effective, secure-by-default cloud platform and Kubernetes ecosystem.
- Design automation, orchestration, observability, and disaster readiness into products and platform services.
- Participate in production support and on-call rotations, providing senior-level guidance during critical events.
- Lead incident management and post-incident retrospectives while coaching teams in reliability practices.
- Mentor colleagues, contribute to architectural discussions, and drive scalable and sustainable technical improvements.
Requirements
- Experience writing design documents and postmortems and refactoring application code.
- Experience building automation to reduce operational burden or developing internal SaaS tools.
- Ability to advocate for and introduce SRE principles such as SLOs and SLAs.
- Experience with public cloud or hosted datacenter environments; Azure and AKS are preferred.
- Passion for collaborative teamwork and influencing reliability best practices across teams.
- Bonus qualifications include Linux server experience, Terraform, Chef, Docker, Prometheus/Grafana or ELK/EFK, CI/CD pipelines, rollout strategies, programming languages such as Java, Python, or Golang, and scripting languages such as PowerShell, Bash, Python, or Ruby.
- A bachelor's degree or equivalent experience in Computer Engineering or a related field is listed as a bonus qualification.
Benefits
- Hybrid schedule with Tuesdays and Thursdays dedicated to in-office collaboration and Mondays and Fridays reserved for remote-friendly focus time.
- Flexible work hours, 25 days of annual leave, additional flexible time off, and company wellness days.
- Annual performance-based bonus, enhanced parental leave, pension matching up to 6%, private medical insurance, healthcare cash plan, group life cover, income protection, and critical illness protection.
- Unlimited access to LinkedIn Learning and Microsoft courses and training, plus opportunities to earn certifications.
- Inclusive and supportive culture, career development framework, modern open-plan workspace, gaming area, free snacks and drinks, and regular social events.
- Access to RethinkCare behavioral health and well-being resources.
Categories
DevOpsSite Reliability
