Rivian

Staff Software Engineer, ML Infra, Autonomy

Rivian
Apply
13 hours ago
Palo Alto, CA, USAStaff+
H1B sponsor

Base Salary

$207k - $258k/yr

Responsibilities

  • Design and evolve a control plane for accelerator clusters across cloud providers and silicon types.
  • Build fleet-scale gang scheduling, multi-tenant resource management, fairness, quota borrowing, preemption, backfill, and automated node and job-failure handling.
  • Design storage tiering, caching, sharding, indexing, and training-native data formats for large-scale random access and sequential scans.
  • Build monitoring, profiling, workload baselines, and optimization tooling across storage, nodes, GPUs, interconnects, schedulers, and individual jobs.
  • Develop a thin training framework with accelerator-specific backends, golden images, validated launch recipes, and model-zoo CI.
  • Coordinate upgrades across distributed compute, Kubernetes operators, queueing, and experiment-tracking systems.
  • Define the platform roadmap, architecture, team boundaries, build-versus-buy decisions, and priorities based on measured goodput and cost.
  • Partner with ML engineers and infrastructure teams to identify platform problems, lead cross-organizational architecture, communicate recommendations, and mentor engineers.
  • Help define the team’s technical direction, operating model, and future hiring.],
  • requirements:[

Requirements

  • At least 5 years of software engineering experience or equivalent demonstrated impact, including substantial production distributed-systems work.
  • Staff-level technical leadership with experience setting strategy, making pragmatic trade-offs, and driving ambiguous initiatives into production.
  • Hands-on experience building or operating large-scale compute or ML infrastructure and diagnosing failures at scale.
  • Deep expertise in at least one of cluster scheduling and multi-tenancy, training data storage and I/O, distributed training frameworks and accelerators, or observability and performance engineering.
  • Strong programming skills in Python and at least one additional relevant language such as Go, Rust, or C++.
  • Experience operating on a major cloud provider, preferably AWS, and on Kubernetes.
  • Strong communication, developer empathy, and a record of building platforms that engineers adopt and trust.
  • Preferred experience includes operating-system internals, large-model training, cross-GPU communication, Ray or KubeRay, Kueue, Kubernetes scheduler extensions, GPU profiling, training-native or columnar data formats, GPU-side video decoding, MLflow, or LLM agents for infrastructure operations.

Benefits

  • Comprehensive benefits are available to eligible full-time and part-time employees, spouses or domestic partners, and children up to age 26.
  • Benefits include paid vacation, paid sick leave, life insurance, medical, dental, vision, short-term disability, and long-term disability insurance.
  • Eligible employees may participate in Rivian’s 401(k) Plan and Employee Stock Purchase Program.
  • Full-time employee coverage begins on the first day of employment; part-time coverage begins the first of the month following 90 days.
Rivian

About Rivian

10,000+ employees

Rivian designs and manufactures electric vehicles, including the R1T pickup, R1S SUV, and commercial delivery vans, selling directly to consumers and fleet operators. Founded in 2009 and headquartered in Irvine, California, it operates a manufacturing plant in Normal, Illinois, and is publicly traded on NASDAQ. The company also develops in-vehicle software and electrical architecture and offers charging and service support for its owners.

Contact me