
ML Engineer, Infrastructure
Prior Labs1 month ago
Berlin, Germany or New York, NY, USAMid Level
Responsibilities
- Own and evolve multi-cluster GPU infrastructure, including architecture, scheduling, reliability, and cost optimization.
- Improve GPU utilization and distributed-training throughput through profiling, memory optimization, and systems-level debugging.
- Architect multi-cluster orchestration, provider diversification, hardware adoption, and compute capacity planning.
- Build developer productivity tooling for CI pipelines, experiment tracking, model registries, data processing, and internal research workflows.
- Manage compute costs and evaluate cost per FLOP across providers and hardware.
- Collaborate directly with researchers and contribute to model training when appropriate.
Requirements
- At least 3 years of experience building and operating production GPU infrastructure or distributed-training systems at scale.
- Deep hands-on experience with Slurm and cluster management, including multi-tenant GPU workload scheduling and utilization optimization.
- Expert systems knowledge covering memory bandwidth, GPU profiling, and hardware-level performance analysis.
- Strong Python skills and genuine fluency with PyTorch internals.
- A track record of improving training throughput or cost efficiency through infrastructure decisions.
- Strong familiarity with AI productivity tooling such as Claude Code, Cursor, or similar tools.
- Preferred experience includes tens-of-millions-scale GPU spending, multi-cloud or hybrid HPC/cloud infrastructure, Triton, CUDA, custom kernels, multi-cluster orchestration, and experiment tracking, model registry, or ML pipeline tooling.
Benefits
- Teams are based in Berlin, Freiburg, and New York, with exceptional remote arrangements generally involving frequent travel to an office.
- The company brings everyone together regularly for offsites.
- The company offers equal opportunities and welcomes applicants from diverse backgrounds.