42dot

Senior AI Data Pipeline Engineer (Autonomous Driving)

42dot
Apply
4 months ago
Sunnyvale, CA, USASenior

Base Salary

$133k - $254k/yr

Responsibilities

  • Develop reliable, high-scale data extraction pipelines that convert fleet-collected raw data into high-value autonomous-driving scene data.
  • Build data-labeling pipelines that perform auto-labeling inferences for autonomous-driving algorithms.
  • Develop data SDKs for scene search, dataset preparation, and dataset loading.
  • Build and maintain the autonomous-driving data lakehouse containing sensor, calibration, and annotation data.
  • Analyze and improve data-processing latency, data-search latency, and test-procedure coverage.
  • Maintain infrastructure for data-processing pipelines, databases, lakehouse systems, and data serving.
  • Collaborate with ML algorithm, ML application, and cloud infrastructure teams on autonomous-driving system architecture.
  • Lead technical projects and align data-platform development with the ML development lifecycle.

Requirements

  • Bachelor's degree or higher in Computer Science, Engineering, Robotics, or a similar technical field.
  • At least 7 years of experience in Data Engineering, DataOps, or ML Platform roles.
  • Proficiency in Python and substantial experience developing Python SDKs.
  • Hands-on experience orchestrating data-pipeline jobs with Databricks Workflows or Apache Airflow and integrating pipelines with machine-learning models.
  • Working experience with databases such as MongoDB and PostgreSQL.
  • Extensive experience with data architectures and technologies such as Hive data warehouses or Delta Lake lakehouses.
  • Experience with Apache Spark or other big-data computing engines.
  • Strong leadership and communication skills, including the ability to lead technical projects.
  • Preferred: experience with autonomous-vehicle sensor data including LiDAR, cameras, or radar.
  • Preferred: experience with ML model-training lifecycles, including data preparation, training, validation, and deployment.
  • Preferred: understanding of PyTorch, TensorFlow, data-governance principles, data-privacy regulations, and data-security implementation.

Tech Stack

Apache AirflowApache HiveApache SparkMongoDBPostgreSQLPythonPyTorchTensorFlow

Categories

Data Engineering
42dot

About 42dot

501-1,000 employees
Contact me