8 hours ago
Base Salary
$210k - $250k/yr
Responsibilities
- Debug and resolve failing or blocked data ingestion pipelines.
- Investigate corrupt, malformed, unexpected, incomplete, or inconsistent data.
- Design resilient ingestion pipelines that prevent individual bad data segments from blocking broader workflows.
- Improve handling of varied data formats from partners, suppliers, and third-party sources.
- Support orchestration of multi-step workflows, including dependencies, retries, and queue management.
- Optimize Apache Spark jobs and data-processing pipelines for throughput, compute efficiency, and reliability.
- Reduce operational toil from failed jobs, stalled pipelines, and manual interventions.
- Operate high-volume batch-processing systems and prioritize datasets for annotation, data science, and model training teams.
- Partner with Data Platform and downstream engineering teams on immediate fixes and scalable long-term solutions.
- Contribute to the technical direction, maintainability, and operational excellence of the ingestion platform.
Requirements
- Strong production experience with Apache Spark and Python.
- Experience building, debugging, or operating large-scale data-ingestion, ETL, or data-processing pipelines.
- Experience with distributed data-processing systems and production pipeline failure debugging.
- Ability to optimize jobs for throughput, compute efficiency, and reliability.
- Experience working with messy, corrupt, incomplete, or inconsistent data.
- Understanding of orchestration across multi-step pipelines and downstream dependencies.
- Experience at significant data scale, ideally petabyte-scale or similarly high-throughput environments.
- Ability to work independently in a fast-moving, highly technical environment with a practical, delivery-focused mindset.
- Desirable experience with Airflow, Flyte, Databricks Workflows, Databricks, Delta Lake, or Delta tables.
- Desirable experience with Scala or Java in Spark-based environments, queue-based processing, retry handling, high-throughput batch processing, or systems with many data producers and consumers.
- Experience with partner or supplier data, cost optimization for compute- and storage-heavy platforms, automotive, robotics, autonomy, mapping, ML data platforms, embodied AI, or relevant distributed-systems work from high-performance domains is advantageous.
Benefits
- Full-time, permanent position based in the Sunnyvale, California office.
- Hybrid working policy combining office and workshop time with working from home.
- Competitive equity package.
- Inclusive workplace and interview accommodations are available upon request.
Tech Stack
Categories
Data Engineering
About Wayve
Wayve builds end-to-end autonomous driving software—the vehicle-agnostic Wayve AI Driver—that runs on onboard compute and native sensors, licensed to automakers and fleet operators. Its platform spans ADAS and higher autonomy (L2+/L3 to robotaxi) and is designed to generalize across vehicle types and geographies. Founded in 2017 and headquartered in London, it tests its models across Europe, North America, and Japan, with a U.S. base in Sunnyvale, CA.
