16 hours ago
Mexico City, MexicoSenior
Responsibilities
- Lead reliability uplift initiatives across assigned applications and services.
- Assess incidents, degradation patterns, operational pain points, manual toil, and reliability risks, then build and execute improvement roadmaps.
- Lead complex production incident reviews, perform root-cause analysis, and drive corrective and preventive actions through completion.
- Improve incident-management workflows, including triage, escalation, response, resolution, communication, and remediation tracking.
- Identify and drive automation for support activities, health checks, data collection, remediation, and reporting.
- Evaluate application resiliency across dependencies, integrations, data flows, error handling, retry logic, timeouts, and recovery mechanisms.
- Identify performance bottlenecks and validate tuning through load testing, profiling, and performance diagnostics.
- Define meaningful health indicators, SLIs, SLOs, and error budgets while improving monitoring, alerting, and end-to-end observability.
- Coordinate cross-functional remediation activities and communicate reliability risks, priorities, metrics, and progress to stakeholders.
Requirements
- 8+ years of experience in application reliability, production engineering, SRE, or application operations.
- Strong diagnostic skills across logs, metrics, traces, and distributed systems.
- Experience with incident management, root-cause analysis, and operational readiness.
- Hands-on experience with monitoring tools such as Splunk, Datadog, New Relic, AppDynamics, and Grafana.
- Strong understanding of application architecture, integrations, APIs, databases, and cloud platforms.
- Experience with automation using Python, Shell, Ansible, Jenkins, or similar tools.
- Excellent communication and stakeholder-management skills.
- Ability to independently lead cross-team reliability initiatives.
Benefits
- Ongoing support and funding for training and development plans.
- Competitive benefits and salary package.
- Opportunity to work for leading global companies.
- Inclusive, non-discriminatory workplace with investment in employee wellbeing and development.
Tech Stack
Categories
Site Reliability
About Cognizant
Cognizant is a public IT services and consulting firm that designs, builds, and runs enterprise technology, including digital engineering, cloud modernization, data/AI, and managed services. It sells consulting, systems integration, and outsourcing on multi-year engagements to large enterprises in healthcare, banking, retail, communications, and manufacturing. Founded in 1994 and headquartered in Teaneck, New Jersey, Cognizant is NASDAQ-listed (CTSH) and a Fortune 500 company.
