2 days ago
Base Salary
$184k - $357k/yr
Responsibilities
- Design and develop learning-based multimodal sensor-fusion systems that combine synchronized sensor history, ego-motion, navigation context, and driving context into unified spatiotemporal world representations.
- Build architectures that reason across camera, LiDAR, radar, and vehicle-state inputs while handling calibration, synchronization, coordinate transforms, latency, and uncertainty.
- Develop end-to-end and multi-task models for road structure, semantic scene understanding, occupancy, free space, and other driving-relevant outputs.
- Develop Transformer-based early, late, and hierarchical fusion architectures using BEV, point/voxel, and image-based representations, temporal aggregation, and cross-modal attention.
- Create training, fine-tuning, and evaluation pipelines for large-scale multimodal datasets and define objectives and metrics for perception, geometry, prediction, latency, and safety.
- Investigate foundation-model approaches for autonomous driving, including vision-language models, multimodal pre-training, representation learning, and efficient learned-world-model deployment.
- Collaborate with perception, mapping, prediction, planning, simulation, data, and embedded-software teams to deliver production-quality AV systems.
- Develop tools to analyze and debug model failures, sensor disagreement, long-tail scenarios, distribution shift, and regressions in simulation and on-road evaluation.
Requirements
- BS, MS, or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a related technical field, or equivalent experience.
- 8+ years of experience, including at least 2+ years in the AV or robotics industry and 2+ years of technical leadership experience.
- Strong experience developing production-quality sensor-fusion, perception, state-estimation, or autonomous-driving systems.
- Experience with learning-based multimodal perception or fusion using at least two of cameras, LiDAR, radar, maps, navigation, and ego-motion signals.
- Strong understanding of 3D geometry, coordinate frames, calibration, temporal synchronization, ego-motion compensation, tracking, uncertainty estimation, and sensor failure modes.
- Experience with deep learning for 3D perception, scene representation, occupancy or occlusion prediction, semantic segmentation, object detection or tracking, motion prediction, or planning.
- Strong C++ and Python skills and hands-on experience developing, training, and optimizing deep-learning models in PyTorch.
- Experience with Transformer, VLM, or multimodal foundation-model architectures, including pre-training, fine-tuning, distillation, quantization, or efficient inference.
- Experience training and evaluating models at scale, including distributed training, dataset curation, offline evaluation, simulation-based validation, and production monitoring.
- Preferred experience with multi-task driving models, BEV or voxel methods, neural scene representation, 3D reconstruction, occupancy-flow, spatiotemporal world models, publications or open-source contributions, and automotive-grade real-time deployment.
Benefits
- Base salary range is $184,000-$287,500 for Level 4 and $224,000-$356,500 for Level 5, with eligibility for equity and benefits.
- Applications will be accepted at least until September 22, 2026.
- This posting is for an existing vacancy.
Categories
ML EngineeringRobotics
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
