
Staff/Principal DevOps Engineer, AI Inference
Lila Sciencesabout 3 hours ago
Cambridge, MA, USAStaff+ / Senior
H1B Sponsor
Base Salary
$192k - $272k/yr
Responsibilities
- Design and implement GPU/accelerator infrastructure on Kubernetes.
- Build model serving platforms using frameworks like vLLM and Triton Inference Server.
- Develop intelligent request routing and load balancing for heterogeneous accelerator fleets.
- Create autoscaling systems to match inference compute supply with demand.
- Establish production-grade deployment pipelines for ML models.
- Utilize infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters.
- Monitor observability and optimize performance for model endpoints.
- Implement CI/CD pipelines for model artifacts and automated inference benchmarking.
- Manage AWS cloud infrastructure for machine learning applications.
- Conduct cost optimization and capacity planning for inference workloads.
Requirements
- Expertise in DevOps, SRE, or Platform Engineering with experience in GPU infrastructure.
- Deep knowledge of Kubernetes for ML workloads, including GPU scheduling and resource management.
- Proficiency in deploying to AWS using infrastructure-as-code tools like Terraform and Helm.
- Experience with model serving infrastructure and request batching.
- Strong understanding of networking for distributed inference.
- Proficiency in Python for automation and integration with ML frameworks.
Benefits
- Comprehensive medical, dental, and vision coverage.
- Employer-paid life and disability insurance.
- Flexible time off with generous company-wide holidays.
- Paid parental leave and educational assistance program.
- Commuter benefits, including bike share memberships.
- Company-subsidized lunch program.