
Senior Site Reliability Engineer
Brevan Howard Investment Management2 months ago
London, United KingdomSenior
Responsibilities
- Architect, deploy, and maintain scalable and reliable infrastructure on Google Cloud Platform using Kubernetes.
- Automate lifecycle and operational workflows with infrastructure-as-code tools, Python, and Bash.
- Own and evolve Terraform-managed cloud resources and Helm-based Kubernetes application deployments.
- Implement monitoring, alerting, and logging for system visibility and proactive issue detection.
- Define and enforce SLOs and SLIs, participate in on-call rotation, and lead post-incident reviews.
- Provide software teams with guidance on deployment strategies, scalability, and cloud-native practices.
- Own infrastructure projects from inception through production operation, including documentation and knowledge transfer.
Requirements
- At least 4 years of hands-on experience with Google Cloud Platform or similar cloud infrastructure.
- Expert-level proficiency managing, scaling, and troubleshooting production Kubernetes environments.
- Deep expertise with Terraform and strong experience with Helm.
- Proficiency in at least one major programming language, preferably Python, plus strong Bash and command-line skills.
- Experience setting up and maintaining modern CI/CD pipelines and implementing monitoring and logging tools.
- Solid understanding of TCP/IP, load balancing, DNS, and Kubernetes cloud-native networking.
- Strong experience with Linux systems.
- Ability to own problems end-to-end, work effectively in a small-team generalist environment, learn new technologies quickly, and communicate clearly with technical and non-technical stakeholders.
- Equivalent experience with AWS, Azure, or similar tools is highly valued.
- Familiarity with Istio, cloud and container security practices, and certifications such as CKAD, CKA, or Professional Cloud DevOps Engineer is preferred.
Tech Stack
Categories
DevOpsSite Reliability