14 days ago
London, United KingdomStaff+
Responsibilities
- Migrate applications from Geneos ITRS, GEM, ELK, Splunk, and AppDynamics to Google Cloud Observability and Grafana.
- Develop, test, and maintain OpenTelemetry Collector or Grafana Alloy configurations for metrics, logs, and traces.
- Create reusable Helm charts and Ansible playbooks for observability-agent deployment across OpenShift, Kubernetes, and virtual machine environments.
- Partner with application teams to instrument applications using OpenTelemetry standards and provide technical onboarding support.
- Design and implement Google Cloud Observability dashboards, log-based metrics, alerts, and SLO/SLI tracking.
- Apply SRE practices including error budgets, toil reduction, resiliency patterns, recovery testing, and chaos engineering.
- Drive automation with Ansible and Terraform to streamline infrastructure deployment and reduce recovery time.
- Troubleshoot complex application, database, network, and operating-system issues in large-scale systems.
- Ensure observability and deployment solutions comply with enterprise standards, security requirements, and regulatory expectations.
Requirements
- Significant professional experience in software development or an equivalent field, with a strong focus on Site Reliability Engineering and observability.
- Hands-on experience with OpenTelemetry or Grafana Alloy collectors.
- Strong understanding of SRE concepts, including SLOs, SLIs, error budgets, and toil reduction.
- Proficiency deploying, managing, and troubleshooting applications on OpenShift and Kubernetes.
- Experience creating, modifying, and managing Helm charts.
- Experience with infrastructure as code, configuration management, and automation tools such as Ansible and Terraform.
- Hands-on experience with observability tools such as Prometheus, Grafana, Loki, Mimir, Tempo, AppDynamics, Splunk, or Google Cloud Observability.
- Experience with public cloud platforms such as Google Cloud Platform and Amazon Web Services.
- Experience migrating applications from legacy monitoring tools such as Geneos ITRS, GEM, ELK, Splunk, and AppDynamics.
- Experience writing or maintaining code in Java, Python, Go, or similar languages.
- Experience delivering software and infrastructure using Agile frameworks.
- Strong communication, collaboration, problem-solving, strategic thinking, and diplomacy skills.
Benefits
- Hybrid working model with up to two days working from home per week.
- 27 days of annual leave plus bank holidays.
- Competitive base salary and eligibility for a discretionary annual performance-related bonus.
- Private medical care and life insurance.
- Employee Assistance Program.
- Pension plan and paid parental leave.
- Special discounts for employees, family, and friends.
- Access to learning and development resources.
Tech Stack
Categories
DevOpsSite Reliability
About Citi
Citi's mission is to serve as a trusted partner to our clients by responsibly providing financial services that enable growth and economic progress. Our core activities are safeguarding assets, lending money, making payments and accessing the capital markets on behalf of our clients. We have over 200 years of experience helping our clients meet the world's toughest challenges and embrace its greatest opportunities. We are Citi, the global bank – an institution connecting millions of people across hundreds of countries and cities. For information on Citi’s commitment to privacy, visit on.citi/privacy.
