1 day ago
Base Salary
$224k - $357k/yr
Responsibilities
- Own the 3D data engine by sourcing, curating, filtering, balancing, and versioning large-scale real-world image and video training corpora.
- Build annotation and auto-labeling pipelines for 3D-grounded supervision, spatial question answering, reference frames, camera motion, geometry, correspondence, and reasoning traces.
- Manage data quality through deduplication, automated scoring, coverage analysis, and balanced pre-training and supervised fine-tuning mixtures.
- Run and validate full model pre-training and supervised fine-tuning pipelines, ensuring reproducibility and diagnosing data, checkpoint, throughput, and loss regressions.
- Build and operate 3D and spatial evaluation suites, benchmarks, dashboards, and traceable evaluation workflows.
- Partner with research scientists to design dataset and ablation experiments and translate research hypotheses into engineering execution.
- Operate multi-node GPU clusters and optimize data throughput, sharding, and dataloader performance.
- Ship capabilities into Cosmos releases and open-source datasets and benchmarks where appropriate.
Requirements
- MS or PhD in Computer Science, Electrical or Computer Engineering, Robotics, or a related field, or equivalent experience.
- 12+ years of experience building deep learning systems in Python with PyTorch or JAX on Linux.
- Deep expertise in 3D computer vision, multi-view geometry, structure-from-motion or SLAM, depth and camera pose estimation, point cloud processing, or 3D reconstruction.
- Hands-on experience with vision-language models and with building training data and evaluations that improve visual grounding and reasoning quality.
- Experience building large-scale multimodal data pipelines for distributed video and image processing, deduplication, captioning, annotation, automated quality metrics, and dataset versioning.
- Experience running and validating large-model training on multi-GPU, multi-node clusters using distributed training and sharding strategies such as data, tensor, and pipeline parallelism or FSDP.
- Ability to design rigorous benchmarks, clean ablations, and evaluations resistant to gaming.
- Strong written and verbal communication and experience partnering with research scientists.
- Preferred: PhD and/or publications at CVPR, ICCV, ECCV, NeurIPS, ICLR, or CoRL in 3D vision, multimodal learning, or embodied AI.
- Preferred: experience with web-scale or petabyte-scale video corpora and distributed processing infrastructure using Ray, Spark, Slurm, or similar tools.
- Preferred: familiarity with 3D foundation models, modern auto-labeling, vision-language model post-training, supervised fine-tuning, chain-of-thought data design, reward modeling, reinforcement learning with verifiable rewards, and evaluation infrastructure such as VLMEvalKit.
Benefits
- Competitive base salary of 224,000 USD to 356,500 USD, plus equity and benefits.
- Generous benefits package.
- Applications accepted at least until September 19, 2026; the posting is for an existing vacancy.
Tech Stack
Categories
About Nvidia
Nvidia designs and sells GPUs and accelerated computing platforms for data centers, AI/ML, graphics, gaming, and automotive, monetizing through hardware, software platforms (CUDA, AI frameworks), and systems like DGX and networking. Customers include cloud providers, enterprises, researchers, and OEMs. Founded in 1993 and headquartered in Santa Clara, it is a public company traded on NASDAQ under NVDA.
