10 days ago
Base Salary
$224k - $280k/yr
Responsibilities
- Design, implement, and scale repeatable Kubernetes-based infrastructure for large-scale distributed GPU training.
- Orchestrate concurrent machine learning workloads across large GPU clusters using distributed compute frameworks.
- Build MLOps, model management, experiment tracking, metadata, and model registry systems.
- Build and optimize high-throughput pipelines for streaming petabyte-scale multi-sensor vehicle logs into training environments.
- Architect autonomous model validation infrastructure and continuous integration testing for regression-free vehicle policy releases.
- Partner with robotics engineers and machine learning researchers to improve training workflows and the deploy-to-vehicle lifecycle.
Requirements
- Have 8+ years of professional software engineering experience.
- Demonstrate strong backend systems programming skills with Go, Python, Java, or similar languages.
- Have Kubernetes expertise and experience building cloud-agnostic environments from scratch.
- Have implemented distributed machine learning compute frameworks such as Ray for large multi-node GPU workloads.
- Have hands-on experience building MLOps pipelines, metadata tracking architectures, and model registries using platforms such as MLflow.
- Have experience managing high-throughput data pipelines with modern distributed data engines.
- Rust familiarity or exposure is a plus.
Benefits
- Medical, dental, vision, disability, and life insurance.
- Flexible Spending Account and Health Savings Account options.
- 401(k) plan and equity eligibility.
- Sick time, unlimited flexible time off, and paid holidays.
- Paid parental leave.
- Pre-tax commuter benefit plan.
- Team lunch in the SoMa office every Tuesday and Thursday.
- This is a full-time exempt role based in San Francisco and requires onsite work five days per week.
