1 hour ago
Hyderābād, IndiaSenior
Responsibilities
- Act as the technical lead for SRE initiatives across multiple product areas.
- Drive SimCorp’s strategic use of Microsoft Azure for onboarding and site reliability.
- Architect scalable, secure, and automated solutions for client onboarding and live operations.
- Lead the evolution of observability, CI/CD pipelines, Infrastructure as Code standards, and disaster recovery frameworks.
- Govern Azure implementation patterns for standardization, reliability, and cost efficiency.
- Solve complex reliability challenges involving distributed cloud systems.
- Advise engineering leads and product owners on cloud platform decisions, trade-offs, and risk mitigation.
- Collaborate with Information Security, Platform Engineering, and Architecture teams on compliance and cloud controls.
- Define SLOs, SLIs, and other reliability metrics across departments.
- Lead root cause analyses, major incident postmortems, and reliability retrospectives.
- Mentor and coach senior and lead engineers and build SRE communities of practice.
- Represent the SRE function in executive planning, roadmap definition, and technical due diligence.
- Contribute to SimCorp’s transformation into a SaaS-first, cloud-native company.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related field.
- 8+ years of experience in Site Reliability Engineering, cloud infrastructure, or platform architecture roles.
- Extensive expertise in Microsoft Azure, including architecture, deployment, automation, and cost optimization.
- Strong understanding of cloud-native and hybrid architectures, distributed systems, networking, and security.
- Mastery of Infrastructure as Code using Terraform, ARM, Bicep, and related tooling.
- Deep knowledge of observability stacks including Azure Monitor, Log Analytics, Grafana, and Application Insights.
- Experience leading complex incident and problem management efforts at scale.
- Broad technical skills including Kubernetes, Docker, CI/CD pipelines, SQL, APIs, and scripting.
- Strong foundation in ITIL processes and operational excellence.
- Ability to influence senior stakeholders, lead through ambiguity, and align engineering with business needs.
- Experience working in or guiding teams in regulated, security-conscious environments such as financial services.
- Demonstrated commitment to mentorship, knowledge sharing, and engineering culture.
- Ability to think strategically while delivering pragmatic, hands-on solutions.
Benefits
- Global hybrid work policy requiring two days per week in the office, with remote work available on other days.
- Every sixth sprint is reserved for planning and innovation to support learning and experimentation.
- High degree of self-direction and autonomy in planning, organizing, and designing work.
- Inclusive and diverse company culture.
- Emphasis on work-life balance.
- Employee empowerment and involvement in shaping work processes.
Tech Stack
Categories
Site Reliability
