1 day ago
Gurgaon, IndiaSenior
Responsibilities
- Own EMS production incidents from investigation and triage through mitigation, recovery, root-cause analysis, and corrective actions.
- Participate in on-call rotations and troubleshoot application, infrastructure, API, integration, latency, capacity, and availability issues.
- Design and maintain real-time operational dashboards and customer-impact-oriented monitoring and alerting.
- Define and monitor SLIs, SLOs, error budgets, availability, latency, throughput, error, and dependency indicators.
- Enhance ELF and automate repetitive support activities, health checks, diagnostics, remediation, reporting, and operational tooling.
- Plan and lead reliability readiness for major events, including capacity planning, load validation, failure simulations, operational rehearsals, and event-day monitoring.
- Improve reliability engineering practices through incident management, postmortems, dependency monitoring, resilience testing, and disaster recovery preparedness.
Requirements
- Strong experience in Site Reliability Engineering, Production Engineering, DevOps, or enterprise-scale production support.
- Strong hands-on production troubleshooting and incident-management experience.
- Experience building production monitoring, alerting, and observability dashboards.
- Strong understanding of application and infrastructure metrics, logs, distributed services, and API monitoring.
- Experience with cloud infrastructure and containerized environments.
- Knowledge of Kubernetes, CI/CD, Infrastructure as Code, and automation.
- Scripting or programming experience with Python, Bash, PowerShell, or equivalent.
- Understanding of networking concepts including DNS, load balancing, firewalls, routing, and connectivity troubleshooting.
- Experience performing root-cause analysis and implementing preventive engineering actions.
- Ability to analyze traffic, latency, errors, and capacity trends and translate findings into engineering improvements.
- Strong communication and coordination skills during high-severity production incidents.
- Preferred experience with Azure, AKS/Kubernetes, Terraform, GitHub Actions, Grafana, Prometheus, Splunk, APM, synthetic monitoring, automated remediation, chaos or game-day testing, and high-volume event readiness.
- Experience supporting highly available, customer-facing platforms during major events and peak traffic periods is strongly preferred.
Tech Stack
Categories
DevOpsSite Reliability
About Cognizant
Cognizant is a public IT services and consulting firm that designs, builds, and runs enterprise technology, including digital engineering, cloud modernization, data/AI, and managed services. It sells consulting, systems integration, and outsourcing on multi-year engagements to large enterprises in healthcare, banking, retail, communications, and manufacturing. Founded in 1994 and headquartered in Teaneck, New Jersey, Cognizant is NASDAQ-listed (CTSH) and a Fortune 500 company.
