1 hour ago
Hyderābād, IndiaSenior
Responsibilities
- Lead SRE initiatives across multiple product areas and serve as a technical authority for reliability strategy.
- Drive strategic use of Microsoft Azure across onboarding and site reliability disciplines.
- Architect scalable, secure, and automated solutions for client onboarding and live operations.
- Design and evolve observability, CI/CD pipeline, infrastructure-as-code, and disaster recovery capabilities.
- Govern Azure implementation patterns for standardization, reliability, and cost efficiency.
- Solve complex reliability challenges involving distributed cloud systems and advise engineering and product leaders on cloud decisions.
- Define reliability metrics such as SLOs and SLIs across departments.
- Lead root cause analyses, major incident postmortems, and reliability retrospectives.
- Mentor senior and lead engineers and build communities of practice around SRE.
- Represent SRE in executive planning, roadmap definition, technical due diligence, and company cloud transformation.
Requirements
- Bachelor’s or master’s degree in computer science, engineering, or a related field.
- 8+ years of experience in Site Reliability Engineering, cloud infrastructure, or platform architecture roles.
- Extensive expertise in Microsoft Azure architecture, deployment, automation, and cost optimization.
- Strong understanding of cloud-native and hybrid architectures, distributed systems, networking, and security.
- Mastery of Infrastructure as Code using Terraform, ARM, Bicep, and related tooling.
- Deep knowledge of observability stacks including Azure Monitor, Log Analytics, Grafana, and Application Insights.
- Experience leading complex incident and problem management efforts at scale.
- Broad technical skills including Kubernetes, Docker, SQL, APIs, and scripting.
- Strong foundation in ITIL processes and operational excellence.
- Ability to influence senior stakeholders, lead through ambiguity, and align engineering with business needs.
- Experience in regulated, security-conscious environments such as financial services.
- Demonstrated commitment to mentorship, knowledge sharing, and engineering culture.
- Ability to combine strategic thinking with pragmatic, hands-on solutions.
Benefits
- Global hybrid work policy with two required office days per week and remote work available on other days.
- Every sixth sprint is reserved for planning and innovation, supporting learning and experimentation.
- High degree of self-direction and autonomy in planning, organizing, and designing work.
- Inclusive and diverse company culture.
- Emphasis on work-life balance and employee well-being.
- Employees are empowered to shape work processes and contribute their perspectives.
Tech Stack
Categories
Site Reliability
