Member of Technical Staff, RL Infra
The Inception Company6 months ago
San Mateo, CA, USAMid Level
Responsibilities
- Design, build, and optimize infrastructure for large-scale reinforcement learning and post-training workloads.
- Improve the reliability, scalability, and throughput of distributed RL training pipelines and workloads.
- Develop monitoring and observability tools to support uptime, debuggability, and reproducibility.
- Collaborate closely with researchers to make RL systems stable, efficient, and production-ready.
Requirements
- BS, MS, or PhD in Computer Science, Engineering, or a related field, or equivalent experience.
- Systems-level understanding of PyTorch, TensorFlow, Ray, and Megatron.
- Experience with reinforcement learning workloads such as PPO, DPO, RLHF, or reward modeling.
- Experience with Docker, Kubernetes, and CI/CD pipelines.
- Preferred experience building and maintaining language models with tens of billions of parameters or more.
- Preferred experience with Kubeflow and Airflow.
- Preferred background in performance optimization and profiling of ML systems.