GrepJob
Lila Sciences

Staff/Principal DevOps Engineer, AI Inference

Lila Sciences
Apply
about 3 hours ago
Cambridge, MA, USAStaff+ / Senior
H1B Sponsor

Base Salary

$192k - $272k/yr

Responsibilities

  • Design and implement GPU/accelerator infrastructure on Kubernetes.
  • Build model serving platforms using frameworks like vLLM and Triton Inference Server.
  • Develop intelligent request routing and load balancing for heterogeneous accelerator fleets.
  • Create autoscaling systems to match inference compute supply with demand.
  • Establish production-grade deployment pipelines for ML models.
  • Utilize infrastructure-as-code with Terraform and Helm for GPU-accelerated EKS clusters.
  • Monitor observability and optimize performance for model endpoints.
  • Implement CI/CD pipelines for model artifacts and automated inference benchmarking.
  • Manage AWS cloud infrastructure for machine learning applications.
  • Conduct cost optimization and capacity planning for inference workloads.

Requirements

  • Expertise in DevOps, SRE, or Platform Engineering with experience in GPU infrastructure.
  • Deep knowledge of Kubernetes for ML workloads, including GPU scheduling and resource management.
  • Proficiency in deploying to AWS using infrastructure-as-code tools like Terraform and Helm.
  • Experience with model serving infrastructure and request batching.
  • Strong understanding of networking for distributed inference.
  • Proficiency in Python for automation and integration with ML frameworks.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • Employer-paid life and disability insurance.
  • Flexible time off with generous company-wide holidays.
  • Paid parental leave and educational assistance program.
  • Commuter benefits, including bike share memberships.
  • Company-subsidized lunch program.