Cognizant

Site Reliability Engineer (SRE) EMS Production & Observability

Cognizant
Apply
1 day ago
Gurgaon, IndiaSenior

Responsibilities

  • Own EMS production incidents from investigation and triage through mitigation, recovery, root-cause analysis, and corrective actions.
  • Participate in on-call rotations and troubleshoot application, infrastructure, API, integration, latency, capacity, and availability issues.
  • Design and maintain real-time operational dashboards and customer-impact-oriented monitoring and alerting.
  • Define and monitor SLIs, SLOs, error budgets, availability, latency, throughput, error, and dependency indicators.
  • Enhance ELF and automate repetitive support activities, health checks, diagnostics, remediation, reporting, and operational tooling.
  • Plan and lead reliability readiness for major events, including capacity planning, load validation, failure simulations, operational rehearsals, and event-day monitoring.
  • Improve reliability engineering practices through incident management, postmortems, dependency monitoring, resilience testing, and disaster recovery preparedness.

Requirements

  • Strong experience in Site Reliability Engineering, Production Engineering, DevOps, or enterprise-scale production support.
  • Strong hands-on production troubleshooting and incident-management experience.
  • Experience building production monitoring, alerting, and observability dashboards.
  • Strong understanding of application and infrastructure metrics, logs, distributed services, and API monitoring.
  • Experience with cloud infrastructure and containerized environments.
  • Knowledge of Kubernetes, CI/CD, Infrastructure as Code, and automation.
  • Scripting or programming experience with Python, Bash, PowerShell, or equivalent.
  • Understanding of networking concepts including DNS, load balancing, firewalls, routing, and connectivity troubleshooting.
  • Experience performing root-cause analysis and implementing preventive engineering actions.
  • Ability to analyze traffic, latency, errors, and capacity trends and translate findings into engineering improvements.
  • Strong communication and coordination skills during high-severity production incidents.
  • Preferred experience with Azure, AKS/Kubernetes, Terraform, GitHub Actions, Grafana, Prometheus, Splunk, APM, synthetic monitoring, automated remediation, chaos or game-day testing, and high-volume event readiness.
  • Experience supporting highly available, customer-facing platforms during major events and peak traffic periods is strongly preferred.

Tech Stack

AzureBashGitHub ActionsGrafanaKubernetesPowerShellPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
Cognizant

About Cognizant

10,000+ employees

Cognizant is a public IT services and consulting firm that designs, builds, and runs enterprise technology, including digital engineering, cloud modernization, data/AI, and managed services. It sells consulting, systems integration, and outsourcing on multi-year engagements to large enterprises in healthcare, banking, retail, communications, and manufacturing. Founded in 1994 and headquartered in Teaneck, New Jersey, Cognizant is NASDAQ-listed (CTSH) and a Fortune 500 company.

Contact me