
Research Platform Engineer
Mistral AI20 hours ago
Palo Alto, CA, USAMid Level
Responsibilities
- Develop services, APIs, controllers, and tooling for ML training, evaluation, fine-tuning, and batch inference.
- Build GPU workload orchestration systems supporting queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
- Provision, allocate, and optimize heterogeneous GPU capacity across clusters and regions.
- Enable multi-cluster workload execution based on capacity, data locality, hardware requirements, and organizational priorities.
- Create self-service workflows for launching, observing, debugging, and reproducing distributed workloads.
- Improve GPU utilization, scheduling latency, startup time, throughput, and infrastructure efficiency.
- Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
- Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.
Requirements
- 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
- Proficiency in Python or Go and experience with production-grade distributed systems.
- Strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
- Understanding of Kueue, Karpenter, Volcano, and Kyverno and the problems they address.
- Understanding of distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
- Familiarity with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
- Understanding of quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
- Ability to diagnose performance and reliability issues across software, orchestration, networking, storage, and hardware.
- Interest in developer experience and creating simple, reliable interfaces for complex infrastructure.
- Ability to work effectively in an ambiguous, fast-moving environment shaped by frontier AI research.
Benefits
- Healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation allowances, and other location-specific benefits may be available depending on country.
- Benefits vary by country and current details are provided on the company's Benefits page.
- The role includes participation in on-call rotations.
Tech Stack
Categories
About Mistral AI
Frontier AI. In your hands. We believe in a future where AI is abundant and accessible. We aspire to empower the world to build with—and benefit from—the most significant technology of our time. Join us: mistral.ai/careers