
Staff/Principal DevOps Engineer, AI Inference
Lila Sciences2 months ago
Base Salary
$192k - $272k/yr
Responsibilities
- Build GPU and accelerator infrastructure on Kubernetes, including scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement.
- Develop model-serving platforms using vLLM, Triton Inference Server, TGI, or custom serving stacks with optimized batching, caching, and request routing.
- Design routing, load balancing, and autoscaling systems across heterogeneous accelerator fleets to maximize utilization and minimize latency.
- Create production deployment pipelines for ML models with canary rollouts, A/B testing, model versioning, safe rollback, and multi-region support.
- Manage GPU-accelerated EKS infrastructure with Terraform and Helm, including node pools, spot/on-demand strategies, and accelerator-specific networking.
- Build observability and performance-optimization systems for GPU utilization, inference latency, token throughput, SLOs, SLIs, and per-request GPU memory profiling.
- Develop CI/CD pipelines for model artifacts, including CUDA-dependent container builds, model registry integration, and automated inference benchmarking.
- Operate AWS ML infrastructure using EKS, accelerated EC2 instances, S3 model storage, EFA networking, and least-privilege IAM.
- Drive accelerator cost optimization, capacity planning, right-sizing, spot strategies, and fleet-wide efficiency reporting.
- Collaborate with ML engineers, research scientists, and software engineers to deliver reliable production inference platforms.
Requirements
- Significant experience in DevOps, SRE, or platform engineering operating GPU or accelerator infrastructure at scale.
- Deep Kubernetes experience for ML workloads, including GPU scheduling, resource quotas, node affinity, and accelerator device management.
- Strong AWS and infrastructure-as-code experience with Terraform and Helm, including GPU-based EKS and EC2 compute.
- Experience with model-serving infrastructure, inference servers, request batching, KV-cache optimization, or LLM-serving frameworks.
- Strong understanding of distributed-inference networking, including high-bandwidth interconnects, NCCL, VPC/PrivateLink, and L4/L7 load balancing.
- Strong proficiency in Python for automation, tooling, and integration with ML frameworks.
- Experience with LLM inference optimization techniques such as continuous batching, speculative decoding, quantization, tensor parallelism, or pipeline parallelism is a bonus.
- Experience with multiple accelerator families, multi-region deployments, Rust or Go, ML-focused SRE practices, model registries, artifact versioning, supply-chain security, or custom observability platforms is preferred.
Benefits
- Full-time U.S. employees receive medical, dental, and vision coverage.
- Benefits include employer-paid life and disability insurance, flexible time off, company-wide holidays, paid parental leave, educational assistance, commuter benefits including bike share memberships for office-based employees, and a subsidized lunch program.
- Full-time employees outside the U.S. receive regionally tailored benefits.
- The role is a full-time position; USD salary ranges apply only to U.S.-based positions, with international salaries set to local market.
Categories
About Lila Sciences
Lila Sciences builds an AI-driven research platform and autonomous lab systems for life science, chemistry, and materials R&D teams. Its products combine large AI models with robotic instruments to plan, run, and analyze experiments, integrating into customer-specific scientific workflows. The privately held company serves pharma, materials, and energy organizations and is venture-backed, including a Series A financing in 2025.