22 hours ago
Sunnyvale, CA, USASenior / Staff+
H1B sponsor
Base Salary
$312k - $389k/yr
Responsibilities
- Shape and execute the reinforcement learning roadmap for driving and core model safety applications.
- Develop and evaluate offline, off-policy, and other reward-guided reinforcement learning methods after behavior cloning.
- Improve reward models and learning signals used to train and evaluate driving policies.
- Build large-scale training and experimentation workflows and diagnose distribution shift, objective misspecification, optimization instability, and evaluation bias.
- Define evidence across offline metrics, open-loop tests, closed-loop simulation, and on-road evaluation.
- Productionize successful methods in the shared machine learning stack and contribute through design reviews, code reviews, and mentoring.
- Lead the technical direction and delivery of a learned emergency trajectory model for safety-critical maneuvers.
Requirements
- Strong track record developing and experimentally validating reinforcement learning or sequential decision-making methods on complex, high-dimensional problems.
- Deep understanding of policy and value learning, off-policy learning, function approximation, distribution shift, and learned-objective failure modes.
- Hands-on experience with behavior cloning, reinforcement learning, or related methods.
- Proficiency in Python and PyTorch, with strong software engineering practices and experience building reliable ML training and evaluation systems.
- Excellent experimental judgment, including forming falsifiable hypotheses, defining useful metrics, running disciplined ablations, and making clear technical decisions.
- Ability to lead a substantial technical area and collaborate effectively across research and engineering boundaries.
- Preferred experience with offline reinforcement learning, imitation learning, reward modeling, preference learning, or post-training of large neural policies.
- Preferred experience in autonomous vehicles, robotics, control, safety-critical physical systems, motion planning, vehicle dynamics, or collision avoidance.
- Preferred experience with closed-loop simulation, off-policy evaluation, uncertainty or calibration, and evaluation under rare or shifted conditions.
- Preferred experience training multimodal, transformer-based, or generative policy models at scale.
- Proficiency in C++, CUDA, distributed training, or production ML performance optimization is desirable.
Benefits
- Full-time role based in the Sunnyvale office.
- Hybrid working policy combining office and workshop time with working from home.
- Competitive equity package.
- Inclusive interview accommodations are available upon request.
Categories
ML EngineeringRobotics
About Wayve
Wayve builds end-to-end autonomous driving software—the vehicle-agnostic Wayve AI Driver—that runs on onboard compute and native sensors, licensed to automakers and fleet operators. Its platform spans ADAS and higher autonomy (L2+/L3 to robotaxi) and is designed to generalize across vehicle types and geographies. Founded in 2017 and headquartered in London, it tests its models across Europe, North America, and Japan, with a U.S. base in Sunnyvale, CA.
