2 months ago
Base Salary
$160k - $241k/yr
Responsibilities
- Contribute to training infrastructure spanning multiple generations of accelerators and multi-cluster scheduling and orchestration.
- Design and operate large-scale batch and streaming data pipelines, including ingestion, storage layout, high-throughput data generation, and storage.
- Design and develop agentic-first ML workflows connecting data, training, and evaluation in introspectable, reproducible, and extensible pipelines.
- Own reliability for critical training and release pipelines through instrumentation, meaningful alerting, on-call practices, and incident response.
- Operate infrastructure supporting distributed GPU training, closed-loop reinforcement learning, workflow orchestration, observability, and cost management.
Requirements
- Bachelor’s, master’s, or doctoral degree in Computer Science, Electrical Engineering, or a closely related field, plus at least one year of relevant work experience.
- Strong proficiency in Python and comfort with C++, Go, or a similar systems language.
- Hands-on experience running production infrastructure on Kubernetes.
- Solid distributed-systems fundamentals and the ability to reason about performance, failure modes, and reliability across complex systems.
- Willingness to work deeply in implementation and raise technical and operational standards.
- Demonstrated ownership of operational maturity through monitoring, alerting, and runbooks.
- Strong working knowledge of GCP is a bonus.
- Experience building large-scale data-generation pipelines is a bonus.
- Experience with Kubernetes-native orchestration for ML workloads is a bonus.
- Depth in GPU and distributed-training internals, including NCCL and collective communication, is a bonus.
- Familiarity with GPU and training observability tooling and diagnosing infrastructure bottlenecks is a bonus.
- Track record of reducing infrastructure cost while improving reliability is a bonus.
Benefits
- Annual performance bonus, equity, and competitive benefits package.
- Base pay range is $160,360-$240,540 depending on level, experience, qualifications, education, location, and skills.
Categories
About Nuro
Nuro builds the Nuro Driver, a Level 4 self-driving platform for robotaxis, commercial fleets, and personally owned vehicles. It licenses its autonomy software and integrates it with automotive-grade hardware for automakers and mobility platforms. Founded in 2016 and headquartered in Mountain View, California, Nuro is privately held and has operated real-world self-driving deployments.
