Research Engineer - AI/RL Infrastructure
Applied Intuitionalmost 2 years ago
Base Salary
$126k - $423k/yr
Responsibilities
- Design and build training and evaluation infrastructure that orchestrates massive GPU clusters over petabytes of multimodal sensor data.
- Build benchmarking, continuous evaluation, and regression-tracking systems for model performance across long-tail real-world driving distributions.
- Develop large-scale data sampling, dataset generation, and data curation pipelines using advanced AI models in a closed-loop data flywheel.
- Enable reliable, efficient, and cost-aware distributed training across heterogeneous cloud environments.
- Collaborate with AI research, autonomy, and platform teams to translate research into production-ready systems.
Requirements
- Experience building and operating production-grade software systems across the full machine learning lifecycle, including training, evaluation, data, and deployment.
- Experience with performance engineering and compute acceleration for large-scale machine learning training, including profiling, bottleneck analysis, and optimization.
- Strong systems-level debugging skills across model code, data pipelines, runtimes, and cluster infrastructure.
- Deep familiarity with the open-source machine learning and systems ecosystem, with judgment about adopting open source versus building in-house.
- Technical experience with PyTorch, CUDA, Ray, Flyte, and Kubernetes.
- Industry experience with relevant topics is preferred, particularly self-driving applications.
- Senior/Staff-level experience and potential Tech Lead or Manager capacity are strongly preferred, though candidates at all experience levels may apply.
Benefits
- Primarily in-office work five days per week, with occasional remote flexibility.
- Health, dental, vision, life, and disability insurance.
- 401(k) retirement benefits with employer match.
- Learning and wellness stipends.
- Paid time off.
- Equity may include options or restricted stock units.