about 3 hours ago
Pune, India or Bengaluru, IndiaSenior / Staff+
H1B Sponsor
Responsibilities
- Design, build, and maintain scalable AWS infrastructure for AI/ML workloads using Terraform.
- Own and evolve GitLab CI/CD pipelines for AI platform services.
- Architect a centralized observability stack using Prometheus and Grafana.
- Define and track DORA metrics to drive delivery and reliability improvements.
- Implement infrastructure security best practices and establish platform governance standards.
- Mentor junior/mid-level engineers through design and code reviews.
Requirements
- 8+ years of experience as a Platform Engineer, Site Reliability Engineer, or DevOps Engineer.
- 3+ years supporting AI/ML or data platform infrastructure with deep AWS expertise.
- Strong hands-on experience with Infrastructure as Code using Terraform.
- Proven expertise in centralized observability using Prometheus and Grafana.
- Solid understanding of DevOps/DORA metrics and their application.
- Experience deploying containerized applications on Kubernetes.
Benefits
- Various health plans.
- Time off plans for vacation and sick time.
- Parental leave options.
- Retirement options.
- Education reimbursement.
- In-office perks.
