
ML Infra Engineer, Data Systems
Physical Intelligence5 days ago
Responsibilities
- Design and build high-throughput pipelines that validate, transform, and featurize raw multimodal data.
- Operate large-scale batch and streaming workflows over massive datasets.
- Design object storage layouts, metadata systems, file formats, and efficient access patterns.
- Build systems for backfills, dataset rebuilds, garbage collection, and large-scale transformations.
- Optimize dataloaders, sharding, prefetching, caching, and throughput for model training.
- Build scalable metadata stores for datasets, annotations, and training artifacts.
- Move petabytes of data efficiently across clusters and environments.
- Implement observability, validation, and guardrails to prevent silent data regressions.
- Collaborate with researchers, engineers, and roboticists to translate evolving data needs into robust systems.
Requirements
- Strong software engineering fundamentals.
- Experience building distributed systems or large-scale data pipelines.
- Ability to reason about performance, memory, I/O, and storage efficiency.
- Familiarity with batch and/or streaming processing systems.
- Experience with object storage systems and data format tradeoffs.
- Ability to design, build, operate, and iterate on systems end-to-end.
- Experience working closely with researchers and unblocking fast-moving projects.
- Preferred: experience with large ML training pipelines or dataloading systems, columnar or custom data formats, ClickHouse, Ray, Flink, Spark, petabyte-scale datasets, or performance debugging in data-heavy systems.
Tech Stack
Categories
Data EngineeringML Engineering