2 months ago
Pune, India or Bengaluru, IndiaStaff+
Responsibilities
- Design, build, and maintain scalable, secure, highly available AWS infrastructure for AI/ML workloads using Terraform.
- Own and evolve GitLab CI/CD pipelines for automated builds, testing, security scanning, and multi-environment deployments.
- Architect observability using Prometheus and Grafana, including dashboards, intelligent alerting, incident response, root-cause analysis, and postmortems.
- Define and track DORA metrics and build self-healing, autoscaling, and cost-optimized Kubernetes infrastructure for AI/ML services.
- Support LLM inference, vector databases, agent frameworks, and production AI/RAG deployments.
- Establish infrastructure security and platform governance standards.
- Partner with AI/ML and data engineering teams and mentor junior and mid-level engineers through design and code reviews.
Requirements
- 8+ years of experience as a Platform Engineer, Site Reliability Engineer, or DevOps Engineer, including 3+ years supporting AI/ML or data platform infrastructure.
- Deep hands-on expertise with AWS services including EKS, Lambda, ECS, VPC, IAM, and S3.
- Strong experience with Terraform infrastructure as code and GitLab CI/CD pipeline design, runners, and deployment automation.
- Expertise with Prometheus, Grafana, Alertmanager, PagerDuty, and OpsGenie for observability, alerting, on-call operations, and incident response.
- Understanding of DORA metrics and their use in improving engineering delivery and reliability.
- Experience operating containerized applications on Kubernetes with Helm, HPA, Cluster Autoscaler, and Karpenter.
- Strong scripting skills in Python and/or Go or Bash.
- Ability to lead complex infrastructure initiatives independently and communicate effectively at staff level.
- Preferred experience with LiteLLM, Ray.io, Anyscale, Temporal, LangGraph, LangChain, LlamaIndex, Qdrant, Pinecone, Weaviate, ZEP, MLflow, Kubeflow, Arize Phoenix, ArgoCD, Flux, Loki, ELK, OpenSearch, Jaeger, and Tempo.
- Preferred knowledge of GitOps, FinOps, and SOC 2 or ISO 27001 security and compliance frameworks.
- Curiosity about AI tools and a history of integrating new technologies into daily workflows.
Benefits
- Various health plans
- Vacation and sick time off plans
- Parental leave options
- Retirement options
- Education reimbursement
- In-office perks
- On-site role based in Bangalore/Pune; the posting also references Zscaler’s hybrid working model
Tech Stack
Categories
DevOpsSite Reliability
About Zscaler
Zscaler builds a cloud-delivered zero trust security platform that replaces traditional network security appliances for enterprises and public-sector organizations. Its Zero Trust Exchange provides secure web gateway, zero trust access, CASB, and cloud firewall services, sold as subscriptions and delivered across 160+ global data centers. Founded in 2007 and headquartered in San Jose, it is a public company on NASDAQ (ZS) serving thousands of customers worldwide.
