Mistral AI

Research Platform Engineer

Mistral AI
Apply
20 hours ago
Palo Alto, CA, USAMid Level

Responsibilities

  • Develop services, APIs, controllers, and tooling for ML training, evaluation, fine-tuning, and batch inference.
  • Build GPU workload orchestration systems supporting queueing, admission control, quotas, priorities, preemption, and topology-aware placement.
  • Provision, allocate, and optimize heterogeneous GPU capacity across clusters and regions.
  • Enable multi-cluster workload execution based on capacity, data locality, hardware requirements, and organizational priorities.
  • Create self-service workflows for launching, observing, debugging, and reproducing distributed workloads.
  • Improve GPU utilization, scheduling latency, startup time, throughput, and infrastructure efficiency.
  • Develop observability, failure recovery, capacity planning, and operational tooling for critical ML workloads.
  • Participate in on-call rotations and troubleshoot issues across applications, schedulers, networking, storage, and GPU infrastructure.

Requirements

  • 4+ years of experience in ML infrastructure, distributed systems, Kubernetes platform engineering, or a related field.
  • Proficiency in Python or Go and experience with production-grade distributed systems.
  • Strong Kubernetes knowledge, including controllers, operators, CRDs, scheduling, networking, storage, and resource management.
  • Understanding of Kueue, Karpenter, Volcano, and Kyverno and the problems they address.
  • Understanding of distributed ML workloads, including training, fine-tuning, evaluation, checkpointing, and batch inference.
  • Familiarity with GPU infrastructure and technologies such as PyTorch, CUDA, NCCL, and high-performance networking.
  • Understanding of quotas, priorities, preemption, gang scheduling, topology awareness, and workload admission.
  • Ability to diagnose performance and reliability issues across software, orchestration, networking, storage, and hardware.
  • Interest in developer experience and creating simple, reliable interfaces for complex infrastructure.
  • Ability to work effectively in an ambiguous, fast-moving environment shaped by frontier AI research.

Benefits

  • Healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal allowances, transportation allowances, and other location-specific benefits may be available depending on country.
  • Benefits vary by country and current details are provided on the company's Benefits page.
  • The role includes participation in on-call rotations.
Mistral AI

About Mistral AI

1,001-5,000 employees

Frontier AI. In your hands. We believe in a future where AI is abundant and accessible. We aspire to empower the world to build with—and benefit from—the most significant technology of our time. Join us: mistral.ai/careers

Contact me