
Machine Learning Infrastructure Engineer
Mind Robotics8 months ago
Palo Alto, CA, USASenior
Responsibilities
- Own distributed training systems and core machine learning infrastructure operating across hundreds of GPUs.
- Improve training-system reliability, performance, and iteration speed.
- Reduce friction in training, evaluating, and deploying models in collaboration with researchers.
Requirements
- Experience building or operating large-scale training systems in PyTorch or JAX.
- Knowledge of sharding, parallelism, and performance optimization for distributed training.
- Ability to work closely with researchers on training, evaluation, and model deployment infrastructure.
Tech Stack
Categories
About Mind Robotics
Mind Robotics builds generalizable industrial robots and the software runtime that lets them perceive, decide, and act under hard real-time constraints. It sells robotic systems and AI-driven middleware for factory-floor tasks, working directly with manufacturers; it has publicly cited a deployment partnership with Rivian. Founded in 2025 and headquartered in Palo Alto, the privately held company focuses on dexterous, adaptive automation for complex production environments.