
Staff Machine Learning Infrastructure Engineer
Dyna Robotics6 months ago
Base Salary
$220k - $320k/yr
Responsibilities
- Architect and own infrastructure for large-scale GPU clusters, including sharding, activation checkpointing, and memory optimization with ZeRO and FSDP.
- Build research codebases and job scheduling systems using Kubernetes and SLURM with fast iteration, automated retries, and failure recovery.
- Design high-throughput pipelines for terabytes of multimodal robot data, including video, proprioception, and 3D signals.
- Build low-latency inference pipelines for real-time robot control using quantization, distillation, and model compilation.
- Profile and optimize GPU utilization, I/O bottlenecks, memory fragmentation, and overall compute-fleet performance.
- Design, build, and operate ML infrastructure end to end to support fast-moving research.
Requirements
- 7+ years of engineering experience with a track record of leading technical projects in high-performance computing or ML infrastructure.
- Deep experience with PyTorch and distributed training frameworks including DeepSpeed and Accelerate.
- Hands-on experience managing cloud GPU environments on GCP or AWS and using Kubernetes for container orchestration.
- Understanding of distributed systems, race conditions, memory management, and NCCL/inter-node communication.
- Experience with mixed precision and gradient accumulation.
- Bonus experience with MCAP, Protobuf, multimodal models, custom Triton kernels, compilers, runtime optimization, or early-stage infrastructure hiring.
Categories
About Dyna Robotics
Dyna Robotics builds general-purpose robotic arms powered by a proprietary embodied‑AI foundation model, sold to businesses to automate repetitive, stationary tasks across manufacturing, logistics, and other settings. Founded in 2024 and headquartered in Redwood City, California, the privately held company reports commercial deployments with customers in multiple industries and is backed by CRV and First Round.