6 days ago
Pune, IndiaStaff+
Responsibilities
- Drive the strategy and implementation of Site Reliability Engineering principles across critical, large-scale applications and services.
- Foster transparency, innovation, accountability, and continuous improvement while communicating SRE initiative progress and impact to stakeholders.
- Ensure applications meet operational resilience requirements and defined impact tolerances.
- Oversee production swing tests, data recovery tests, disaster recovery planning, resiliency testing, and chaos engineering practices.
- Develop automation such as One Touch Recovery solutions to reduce recovery time.
- Partner with development teams to apply cloud-native services and resiliency patterns for reliable, scalable, fault-tolerant applications.
- Develop and scale observability solutions for metrics, logging, and tracing, and help teams instrument applications to monitor system health and performance.
- Analyze complex application, database, network, and operating-system issues in large-scale customer-facing systems.
- Present technical strategy to senior and executive-level audiences and collaborate across business and technical teams.
Requirements
- 13 or more years of deep understanding of SRE concepts, including SLOs, SLIs, error budgets, and toil reduction.
- Significant professional experience in production management, software development, or an equivalent field with a strong focus on Site Reliability Engineering.
- Demonstrated experience with disaster recovery planning, resiliency testing, and fault-tolerant distributed-system design.
- Proficiency deploying, managing, and troubleshooting applications on OpenShift and Kubernetes.
- Hands-on experience with Prometheus, Grafana, Loki, Mimir, Tempo, or AppDynamics.
- Experience with Infrastructure as Code, configuration management, and automation tools such as Ansible and Terraform.
- Experience creating, modifying, and managing Helm charts for application deployment.
- Expertise analyzing complex application, database, network, and operating-system issues in large-scale customer-facing systems.
- Experience with major public cloud providers such as Google Cloud, AWS, or Azure is desired.
- Experience delivering software and infrastructure using Agile frameworks is desired.
- Experience writing or maintaining code in Java, Python, Go, or similar languages is desired.
- Strong communication, diplomacy, problem-solving, strategic-thinking, and cross-functional collaboration skills.
Categories
Site Reliability
About Citi
Citi's mission is to serve as a trusted partner to our clients by responsibly providing financial services that enable growth and economic progress. Our core activities are safeguarding assets, lending money, making payments and accessing the capital markets on behalf of our clients. We have over 200 years of experience helping our clients meet the world's toughest challenges and embrace its greatest opportunities. We are Citi, the global bank – an institution connecting millions of people across hundreds of countries and cities. For information on Citi’s commitment to privacy, visit on.citi/privacy.
