4 days ago
Base Salary
$207k - $258k/yr
Responsibilities
- Own the architecture and roadmap for Autonomy’s distributed compute platform, including scheduling, quotas, fair-share, autoscaling, spot capacity, and preemption across multiple clouds.
- Build scalable tools and APIs that translate high-level job requests into executed distributed work.
- Profile workloads, identify bottlenecks, improve cluster efficiency, eliminate GPU fragmentation, and optimize capacity utilization.
- Own petabyte-scale storage architecture, including layout, partitioning, tiering, lifecycle policies, caching, and prefetching.
- Own multi-cloud compute and data paths, including replication, consistency, and cross-provider data-transfer economics.
- Improve throughput and turnaround time for training-data loading, log replay, and large-scale batch resimulation.
- Operate queue-based and event-driven infrastructure for job submission and autoscaling.
- Define platform metrics, dashboards, alerts, SLAs, on-call strategy, and reliability practices.
- Manage compute and storage costs across AWS accounts and services with per-team and per-workload instrumentation.
- Collaborate with security and privacy teams on governance, access control, retention, and auditing of vehicle data.
- Set technical standards and provide design reviews, code reviews, mentorship, and technical leadership.
Requirements
- Bachelor’s degree in Computer Science, Electrical Engineering, or a related field, or equivalent experience.
- 8+ years of software engineering experience building and operating distributed systems in production.
- 5+ years owning distributed compute, batch execution, or scheduling platforms used by other engineering teams.
- 5+ years of Kubernetes experience involving scheduling, resource management, custom controllers or operators, autoscaling, and large-scale failure modes.
- 5+ years of experience with large-scale object storage, including layout, partitioning, lifecycle management, tiering, caching, performance, and cost tradeoffs.
- 3+ years with Ray, Spark, Flink, or Dask.
- 3+ years of hands-on experience with Go, C++, or Rust, plus strong Python skills.
- 3+ years with Terraform, AWS CDK, or CloudFormation, infrastructure as code, configuration management, and container build and supply chain.
- 3+ years with SQS, Kafka, Kinesis, or equivalent queueing and event-driven systems.
- 3+ years defining production monitoring and alerting with Prometheus, Grafana, Datadog, or CloudWatch.
- 3+ years debugging production distributed systems and conducting incident root-cause analysis.
- 2+ years carrying and meeting an infrastructure cost target.
- Ability to turn ambiguous requirements into detailed system designs and drive implementation independently.
- Demonstrated technical influence through designs, standards, and mentorship.
- GPU cluster management, topology-aware placement, gang scheduling, multi-tenant sharing, and fragmentation experience is preferred.
- Autonomous vehicle, robotics, or continuous high-volume sensor-data experience is preferred.
- Multi-cloud, hybrid-cloud, on-premises, and inter-provider data-movement experience is preferred.
- Linux internals experience, including CPU scheduling, memory management, file systems, and networking, is preferred.
Benefits
- Comprehensive benefits are available to full-time and part-time employees and eligible dependents, including paid vacation, paid sick leave, life insurance, medical, dental, vision, short-term disability, and long-term disability coverage.
- Eligible employees may participate in Rivian’s 401(k) Plan and Employee Stock Purchase Program.
- Full-time employee coverage begins on the first day of employment; part-time coverage begins on the first day of the month after 90 days.
- The role is based in the San Francisco Bay Area for the disclosed salary range.
Tech Stack
Apache FlinkApache KafkaApache SparkAWSC++DatadogGoGrafanaKubernetesLinuxPrometheusPythonRustTerraform
Categories
About Rivian
Rivian designs and manufactures electric vehicles, including the R1T pickup, R1S SUV, and commercial delivery vans, selling directly to consumers and fleet operators. Founded in 2009 and headquartered in Irvine, California, it operates a manufacturing plant in Normal, Illinois, and is publicly traded on NASDAQ. The company also develops in-vehicle software and electrical architecture and offers charging and service support for its owners.
