Lila Sciences

Staff/Principal DevOps Engineer, AI Inference

Lila Sciences
Apply
2 months ago
Cambridge, MA, USAStaff+
H1B sponsor

Base Salary

$192k - $272k/yr

Responsibilities

  • Build GPU and accelerator infrastructure on Kubernetes, including scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement.
  • Develop model-serving platforms using vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing.
  • Design routing, load balancing, and autoscaling systems across heterogeneous accelerator fleets to maximize utilization and minimize latency.
  • Create production deployment pipelines for ML models with canary rollouts, A/B testing, model versioning, safe rollback, and multi-region support.
  • Manage GPU-accelerated EKS infrastructure with Terraform and Helm, including node pools, spot/on-demand strategies, and accelerator-specific networking.
  • Build observability and performance-optimization systems for GPU utilization, inference latency, token throughput, SLOs, SLIs, and per-request GPU memory profiling.
  • Develop CI/CD pipelines for model artifacts, including CUDA-dependent container builds, model registry integration, and automated inference benchmarking.
  • Operate AWS ML infrastructure using EKS, accelerated EC2 instances, S3 model storage, EFA networking, and least-privilege IAM.
  • Drive accelerator cost optimization, capacity planning, right-sizing, spot strategies, and fleet-wide efficiency reporting.
  • Collaborate with ML engineers, research scientists, and software engineers to deliver reliable production inference platforms.

Requirements

  • Significant experience in DevOps, SRE, or platform engineering operating GPU or accelerator infrastructure at scale.
  • Deep Kubernetes experience for ML workloads, including GPU scheduling, resource quotas, node affinity, and accelerator device management.
  • Strong AWS and infrastructure-as-code experience with Terraform and Helm, including GPU-based EKS and EC2 compute.
  • Experience with model-serving infrastructure, inference servers, request batching, KV-cache optimization, or LLM-serving frameworks.
  • Strong understanding of distributed-inference networking, including high-bandwidth interconnects, NCCL, VPC/PrivateLink, and L4/L7 load balancing.
  • Strong proficiency in Python for automation, tooling, and integration with ML frameworks.
  • Experience with LLM inference optimization techniques such as continuous batching, speculative decoding, quantization, tensor parallelism, or pipeline parallelism is a bonus.
  • Experience with multiple accelerator families, multi-region deployments, Rust or Go, ML-focused SRE practices, model registries, artifact versioning, supply-chain security, or custom observability platforms is preferred.

Benefits

  • Full-time U.S. employees receive medical, dental, and vision coverage.
  • Benefits include employer-paid life and disability insurance, flexible time off, company-wide holidays, paid parental leave, educational assistance, commuter benefits including bike share memberships for office-based employees, and a subsidized lunch program.
  • Full-time employees outside the U.S. receive regionally tailored benefits.
  • The role is a full-time position; USD salary ranges apply only to U.S.-based positions, with international salaries set to local market.
Lila Sciences

About Lila Sciences

501-1,000 employees

Lila Sciences builds an AI-driven research platform and autonomous lab systems for life science, chemistry, and materials R&D teams. Its products combine large AI models with robotic instruments to plan, run, and analyze experiments, integrating into customer-specific scientific workflows. The privately held company serves pharma, materials, and energy organizations and is venture-backed, including a Series A financing in 2025.

Contact me