Cognizant

Senior Site Reliability Engineer

Cognizant
Apply
7 days ago
Arizona City, AZ, USASenior
H1B sponsor

Responsibilities

  • Own reliability, availability, performance, compliance, and operational accountability for AI, LLM, and ML-enabled payer platforms.
  • Lead operational governance, continuous improvement, incident management, root cause analysis, problem management, and service restoration.
  • Oversee MLOps processes including model deployment, monitoring, retraining, validation, and rollback.
  • Deploy and manage containerized services with Kubernetes and provision cloud and on-premises infrastructure with Terraform.
  • Standardize configuration management using Ansible and support Node.js services and APIs integrating AI and machine learning capabilities.
  • Establish Git-based version control, release management, code review, and repository governance practices.
  • Drive Linux security hardening, patch management, performance optimization, monitoring, and production triage.
  • Guide Python development for data pipelines, AI orchestration, automation frameworks, and analytics workloads.
  • Monitor platform health using metrics, logs, traces, Prometheus, Grafana, and other observability tools.
  • Collaborate with stakeholders, product owners, platform engineers, AI specialists, operations teams, and healthcare-domain experts.

Requirements

  • Strong experience with Kubernetes and Google Cloud Platform, including Google Kubernetes Engine (GKE).
  • Strong experience with Terraform, Helm, and GitHub Actions for infrastructure and delivery automation.
  • Proficiency in Python, Ansible, and Node.js.
  • Strong experience with Prometheus and Grafana observability tooling.
  • Solid understanding of Linux systems and networking fundamentals.
  • Experience with incident management, on-call support, production triage, automation, and CI/CD pipelines.
  • Strong understanding of AI/ML concepts and AIOps practices, including model lifecycle management, monitoring, or AI-driven alerting.
  • Google Cloud Architect Certification and Certified Kubernetes Administrator certification are helpful.
  • Experience with Java/J2EE, Spring Boot, ML/AI platforms or pipelines, AIOps tools, anomaly detection, predictive analytics, distributed systems, microservices, GPU workloads, Kubeflow, Vertex AI, ML pipelines, or AI-driven monitoring automation is helpful.

Benefits

  • Medical, dental, vision, and life insurance, subject to eligibility.
  • Paid holidays and paid time off.
  • 401(k) plan and contributions.
  • Long-term and short-term disability coverage.
  • Paid parental leave.
  • Employee Stock Purchase Plan.
  • Hybrid work arrangement requiring attendance at a client or Cognizant office based on project needs.

Tech Stack

AnsibleGitHub ActionsGoogle Cloud PlatformGrafanaHelmJavaKubernetesLinuxNode.jsPrometheusPythonSpring BootTerraform

Categories

DevOpsSite Reliability
Cognizant

About Cognizant

10,000+ employees

Cognizant (Nasdaq: CTSH) is an AI Builder and technology services provider, bridging the gap between AI investment and enterprise value. We build full-stack AI solutions powered by deep industry, process and engineering expertise — embedding an organization's unique context into technology systems that amplify human potential and drive tangible outcomes. From strategy to deployment, we help global enterprises move from AI ambition to AI impact and stay ahead in a fast-changing world. See how at cognizant.ai | Follow us @cognizant

Contact me