
Site Reliability Engineer
Cantor Fitzgerald / BGC Partners2 months ago
Belfast, United KingdomSenior
Responsibilities
- Design, build, and maintain highly available, scalable, and resilient production infrastructure.
- Implement Infrastructure as Code solutions for automated provisioning and configuration.
- Enhance CI/CD pipelines and manage containerized workloads across orchestration platforms.
- Build monitoring, alerting, and observability solutions for rapid issue resolution.
- Automate operational workflows, deployments, and administrative tasks.
- Troubleshoot complex production incidents and participate in incident management and root cause analysis.
- Collaborate with development teams to improve system reliability and performance.
- Contribute to disaster recovery and platform resilience strategies.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Engineering.
- Hands-on expertise with Infrastructure as Code technologies such as Terraform and Ansible.
- Experience with Kubernetes, Nomad, or OpenShift.
- Proven experience designing and maintaining CI/CD pipelines.
- Strong Linux systems administration and Bash, Shell, and Python scripting skills.
- Experience with Grafana, InfluxDB, and automation.
- Ability to troubleshoot complex distributed systems and applications.
- Familiarity with Git and modern source control workflows.
- Understanding of high availability, fault tolerance, and disaster recovery principles.
- Strong communication and collaboration skills.
Tech Stack
Categories
DevOpsSite Reliability