6 hours ago
Base Salary
$162k - $234k/yr
Responsibilities
- Design canonical schemas and representations for autonomous-driving drives, scenarios, clips, frames, trajectories, sensor references, vehicle state, map context, labels, predictions, and dataset manifests.
- Build and maintain validated, versioned datasets at more than 10 million records or samples on local or cloud object storage.
- Evaluate and optimize Parquet, Apache Arrow, Lance, partitioning, indexing, file sizing, compaction, schema evolution, lineage, reproducibility, filtering, joins, sampling, and retrieval.
- Build reliable batch or distributed pipelines for ingestion, transformation, enrichment, validation, cataloging, publication, retries, backfills, idempotency, and observability.
- Create self-service dataset discovery and composition for autonomy scenarios, hard examples, balanced datasets, and leakage-resistant train, validation, and test splits.
- Define quality gates, dashboards, alerts, statistical monitoring, drift and anomaly detection, out-of-distribution detection, and root-cause analysis.
- Design model-in-the-loop and human-in-the-loop labeling workflows, including confidence thresholds, review routing, audit sampling, disagreement handling, and label provenance.
- Close the loop from model failures and edge cases through selection, annotation, review, dataset publication, training, and evaluation.
- Prepare replayable trajectory datasets for imitation learning and offline reinforcement learning.
- Own dataset architecture, storage and query performance, data quality, reliability, reproducibility, and cost as a staff-level escalation point.
- Align vehicle collection, autonomy metadata, model interfaces, annotation, storage, governance, and training requirements across partner teams.
- Mentor engineers and strengthen design reviews, testing, observability, reproducibility, cost awareness, data-engineering standards, and operating procedures.
Requirements
- At least 8 years of experience in applied machine learning or deep learning.
- At least 4 years of experience in reinforcement learning, computer vision, or autonomous-driving/ADAS systems.
- Master’s degree in Computer Science, Robotics, Electrical Engineering, Applied Mathematics, or a related field; equivalent industry experience may be considered.
- Production ownership of large data pipelines and datasets exceeding 10 million records or samples, or comparable multi-terabyte scale.
- Deep experience with dataset schema design, columnar storage, partitioning, compression, predicate pushdown, projection, indexing, compaction, schema evolution, versioning, lineage, and reproducibility.
- Hands-on experience with Parquet, Lance, Apache Arrow, or a comparable analytical dataset technology.
- Strong Python and SQL skills with production software-engineering practices, including testing, code review, version control, profiling, debugging, maintainable APIs, and operational documentation.
- Experience building distributed or parallel data pipelines and orchestration using Ray, Spark, Beam, Dask, Airflow, Dagster, or comparable systems.
- Experience with local or cloud object storage such as S3 or GCS, including throughput, caching, data movement, reliability, access control, and cost considerations.
- Knowledge of multimodal autonomous-driving or robotics data, including sensors, synchronization, calibration, ego state, coordinate frames, trajectories, scenarios, and annotations.
- Experience with auto-labeling, active learning, model-in-the-loop or human-in-the-loop systems, annotation QA, statistical monitoring, distribution comparison, drift measurement, anomaly detection, or out-of-distribution detection.
- Experience optimizing distributed ML data loading, caching, sampling, preprocessing, and the data-to-GPU path.
- Experience preparing sequential or trajectory data for imitation learning, offline reinforcement learning, preference or ranking data, reward modeling, or policy evaluation.
- Familiarity with data governance, privacy, retention, redaction, access control, auditability, and security requirements for fleet-collected sensor data.
- A PhD in a relevant field is desired.
Benefits
- Medical, dental, and vision insurance.
- 401(k) with employer match and a defined contribution plan.
- Short- and long-term disability, basic life, and AD&D insurance.
- Employee assistance program, tuition reimbursement, and student loan repayment plans.
- Maternity and non-primary caregiver leave and adoption assistance.
- Employee referral program, vacation, and paid holidays.
- Vehicle lease program covering registration and insurance fees.
- The role is based in Mountain View, California, with a stated salary range of $161,710–$234,325; performance-based merit increases and an annual bonus are also offered.
Tech Stack
Categories
Data EngineeringML Engineering
About Cariad
CARIAD builds the software platforms and digital functions that power vehicles across the Volkswagen Group, including Volkswagen, Audi, and Porsche. Its portfolio spans advanced driver assistance, infotainment and connectivity, charging and driving features, plus backend and cloud services delivered to group brands rather than direct retail. Founded in 2020 and headquartered in Wolfsburg, Germany, CARIAD operates as the automotive software unit of Volkswagen Group.
