Physical Intelligence

ML Infra Engineer, Data Systems

Physical Intelligence
Apply
5 days ago

Responsibilities

  • Design and build high-throughput pipelines that validate, transform, and featurize raw multimodal data.
  • Operate large-scale batch and streaming workflows over massive datasets.
  • Design object storage layouts, metadata systems, file formats, and efficient access patterns.
  • Build systems for backfills, dataset rebuilds, garbage collection, and large-scale transformations.
  • Optimize dataloaders, sharding, prefetching, caching, and throughput for model training.
  • Build scalable metadata stores for datasets, annotations, and training artifacts.
  • Move petabytes of data efficiently across clusters and environments.
  • Implement observability, validation, and guardrails to prevent silent data regressions.
  • Collaborate with researchers, engineers, and roboticists to translate evolving data needs into robust systems.

Requirements

  • Strong software engineering fundamentals.
  • Experience building distributed systems or large-scale data pipelines.
  • Ability to reason about performance, memory, I/O, and storage efficiency.
  • Familiarity with batch and/or streaming processing systems.
  • Experience with object storage systems and data format tradeoffs.
  • Ability to design, build, operate, and iterate on systems end-to-end.
  • Experience working closely with researchers and unblocking fast-moving projects.
  • Preferred: experience with large ML training pipelines or dataloading systems, columnar or custom data formats, ClickHouse, Ray, Flink, Spark, petabyte-scale datasets, or performance debugging in data-heavy systems.

Tech Stack

Apache FlinkApache SparkClickHouse

Categories

Data EngineeringML Engineering
Physical Intelligence

About Physical Intelligence

201-500 employees
Contact me