Eka

Machine Learning / Reinforcement Learning Infrastructure Engineer

Eka
Apply
over 1 year ago
Boston, MA, USASenior

Responsibilities

  • Design, implement, and maintain large-scale model training infrastructure for job orchestration, scheduling, checkpointing, and experiment tracking.
  • Build intuitive tooling for launching, monitoring, debugging, and reproducing research experiments.
  • Scale reinforcement learning and machine learning pipelines across distributed compute clusters.
  • Manage the efficient allocation and utilization of cloud-based compute resources.
  • Partner with researchers to develop scalable training infrastructure, guide training-at-scale best practices, and contribute to core JAX model and training code.
  • Build automated testing pipelines, CI/CD workflows for machine learning, and custom logging and telemetry systems.

Requirements

  • Bachelor’s degree or higher in Computer Science, Computer Engineering, Machine Learning, or a related technical field.
  • Strong software engineering fundamentals and a proven track record building ML training infrastructure, internal developer platforms, or scalable systems.
  • Hands-on experience with large-scale training using JAX, PyTorch, or TensorFlow.
  • Familiarity with distributed training, multi-host systems, data pipelines, and workloads on cloud platforms or orchestration systems such as Kubernetes, SLURM, GCP, or AWS.
  • Strong cross-functional communication, ownership, and developer-experience orientation.
  • Preferred experience in robotics, reinforcement learning, or other machine learning systems.
  • Preferred experience designing abstractions that balance researcher flexibility with system reliability.
Eka

About Eka

11-50 employees
Contact me