
Site Reliability Engineer - Mumbai - India
Plante Moran13 days ago
Mumbai, IndiaMid Level
Responsibilities
- Serve as the Site Reliability Engineer for Application Development, championing reliability, observability, performance, and operational excellence across .NET and React applications.
- Define and track SLOs, SLIs, and error budgets and partner with development teams on reliability improvements.
- Design and maintain monitoring, alerting, logging, and observability solutions using tools such as Azure Monitor, Application Insights, Log Analytics, and Grafana.
- Lead incident response, root cause analysis, blameless postmortems, corrective actions, and production follow-through.
- Use GitHub Copilot, Copilot in Azure, and similar AI tools for incident triage and root cause analysis.
- Build AI-driven internal tooling such as ChatOps bots, runbook copilots, and log-analysis assistants.
- Partner with Application Development, Infrastructure, Security, and Networking teams on resilient, scalable, secure full-stack solutions.
- Write SQL queries for investigation, analysis, and performance tuning.
- Drive system, performance, chaos, and resilience testing.
- Mentor developers on reliability, observability, and production readiness and maintain SRE documentation.
- Participate in continuous improvement initiatives and an on-call rotation.
Requirements
- Bachelor’s degree in computer science, information technology, or a related field.
- At least 3 years of experience troubleshooting and debugging applications built with C# .NET, React, MS SQL, and APIs.
- Experience as a full-stack developer with progressive responsibility in SRE and DevOps responsibilities.
- Hands-on experience with Azure App Service, Functions, Storage, Queues, Key Vault, API Management, Application Insights, and Log Analytics.
- Experience with Git source control, GitHub, Azure DevOps, CI/CD pipelines, feature-flag releases, and GitHub Actions.
- Understanding of blue/green, canary, and feature-flag-based release patterns.
- Experience reviewing and analyzing infrastructure as code using Bicep, ARM, or Terraform.
- Strong understanding of SRE principles, including SLOs, SLIs, error budgets, toil reduction, incident management, and blameless postmortems.
- Strong coding, debugging, and troubleshooting skills with knowledge of relational databases, data models, and distributed systems.
- Experience using Postman or a similar tool to test APIs and conducting code reviews through pull requests.
- Experience with AIOps platforms or AI-driven observability is an added advantage.
- Strong leadership, mentoring, analytical, problem-solving, communication, and interpersonal skills.
- Ability to translate technical issues for non-technical audiences, work in a fast-paced environment, and participate in on-call support.
Benefits
- Plante Moran promotes career growth, employee well-being, professional development, and a supportive workplace culture.
- The role follows a “Workplace for Your Day” model with a principally in-person approach and flexibility determined with the supervisor and team.
- The company provides an inclusive, diverse, and equitable workplace and is an Equal Opportunity Employer.
- The position includes participation in an on-call rotation.