Waabi

Senior / Staff ML Ops Engineer

Waabi
Apply
5 days ago
Remote, United States +6 moreSenior / Staff+

Base Salary

$184k - $272k/yr

Responsibilities

  • Build and evolve Kubernetes-based ML training infrastructure, including GPU scheduling, autoscaling, distributed jobs, capacity strategy, operators, and workflow engines.
  • Develop developer-facing CLIs, SDKs, APIs, job submission workflows, templates, and paved paths for ML teams.
  • Improve training iteration speed, local development, failure diagnosis, and platform reliability through measurement and tooling.
  • Evaluate and adopt ML infrastructure frameworks and tooling through prototypes, migration plans, and user collaboration.
  • Build dataset versioning, sharding, high-throughput multimodal sensor-data loading, experiment tracking, model registry, and lineage capabilities.
  • Create durable Python libraries, CLIs, services, documentation, dashboards, and observability for utilization, throughput, failures, queue times, and experiment cost.
  • Ship CI/CD for models alongside autonomy and simulation systems, and build guardrails for access control, data handling, and cost governance.
  • Drive adoption through prototyping with users, onboarding, office hours, documentation, and support.

Requirements

  • 5+ years of software or infrastructure engineering experience, including platforms used by other engineers and ML or data-intensive production systems.
  • Hands-on Kubernetes expertise covering GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and debugging clusters under load.
  • Excellent Python skills and experience designing APIs and user-friendly CLIs.
  • Practical AWS experience with object storage, IAM, GPU compute, networking, cost management, and infrastructure as code using Terraform, Pulumi, or similar.
  • Experience with distributed training in PyTorch, including DDP, FSDP, or similar, plus experiment tracking and model registry tooling.
  • Fluency with containers, CI/CD, modern build systems, and large monorepos.
  • Ability to influence without authority, evaluate frameworks, pilot solutions, and persuade senior engineers to adopt changes.
  • Strong collaboration, user empathy, product instincts, writing skills, and autonomy in ambiguous environments.
  • Passion for self-driving technologies, frontier AI, and small high-performing teams.
  • Preferred experience with internal developer or research platforms, large-scale distributed GPU training, NCCL, high-performance cluster networking, LiDAR or camera data, Parquet, WebDataset, workflow systems, build systems, remote caching, simulation infrastructure, ML, robotics, autonomous-systems infrastructure, security-sensitive environments, or open-source ML infrastructure contributions.

Benefits

  • Competitive compensation and equity awards.
  • Medical, dental, and vision coverage for full-time employees.
  • Unlimited vacation.
  • Flexible hours and work-from-home support.
  • Daily drinks, snacks, and catered meals when in the office.
  • Regular on-site, off-site, and virtual team-building activities and social events.
Waabi

About Waabi

201-500 employees

Waabi builds AI-driven self-driving technology for long‑haul freight, centered on the Waabi Driver autonomous trucking system and the Waabi World simulator. The company partners with carriers and vehicle makers to commercialize autonomous trucks and explores robotaxi applications. Founded in 2021 by Raquel Urtasun, Waabi is privately held, headquartered in Toronto, with offices in San Francisco, Dallas, and Pittsburgh.

Contact me