6 months ago
Edinburgh, United Kingdom or London, United KingdomStaff+
Responsibilities
- Maintain high availability, scalability, and performance of FNZ platforms.
- Implement monitoring, alerting, and observability solutions to detect and resolve issues proactively.
- Design and implement deployment pipelines and integrate applications with infrastructure and network components.
- Provision and manage infrastructure across on-premises and cloud environments using Terraform.
- Operate and optimize workloads on-premises and in public-cloud environments.
- Manage and troubleshoot application delivery networks, load balancing, and traffic routing.
- Configure and support F5 Distributed Cloud or similar CDN/ADC technologies.
- Participate in on-call rotations, conduct root cause analysis, and implement preventive measures.
- Collaborate with Application, Infrastructure, and Network Engineering teams to deliver reliable services.
Requirements
- Deep understanding of Kubernetes and cluster management.
- Strong experience with Terraform for infrastructure as code across cloud and on-premises environments.
- Hands-on experience with at least one public-cloud platform: AWS, Azure, or GCP.
- Knowledge of F5 Distributed Cloud or similar CDN/ADC platforms and their integration.
- Expertise in application delivery networks, load balancing, traffic routing, and troubleshooting.
- Familiarity with observability tools such as Splunk or New Relic.
- Proficiency with Terraform, Bash, or similar scripting and automation languages.
- Experience with CI/CD pipelines and GitOps workflows is desirable.
- Knowledge of SRE principles and security best practices is desirable.
- Strong problem-solving, troubleshooting, collaboration, and automation skills.
Tech Stack
Categories
Site Reliability