Reflection

Member of Technical Staff - Pre-Training Infra

Reflection
Apply
5 months ago
London, United Kingdom +2 moreSenior
H1B Sponsor

Responsibilities

  • Build and scale distributed training systems for frontier model pre-training.
  • Design and operate large-scale training runs with research teams.
  • Develop infrastructure for efficient training across thousands of GPUs.
  • Optimize training throughput, stability, communication, memory usage, and GPU utilization.
  • Build and maintain pipelines for large-scale datasets, checkpointing, and experiment iteration.
  • Debug performance bottlenecks involving model parallelism, GPU communication, and training runtime systems.
  • Develop systems that enable rapid experimentation with new training techniques.

Requirements

  • Experience building or operating distributed training systems for large machine learning models.
  • Strong experience with distributed training frameworks such as Megatron and DeepSpeed or similar systems.
  • Familiarity with data, tensor, pipeline, or expert parallelism strategies.
  • Experience optimizing training throughput and GPU utilization in large distributed environments.
  • Familiarity with NCCL and performance tuning for distributed workloads.
  • Experience working with ML researchers to productionize experimental training workflows.
  • Strong debugging skills across GPU compute, distributed training systems, and large-scale ML pipelines.
  • Experience working with large datasets and training pipelines for foundation-model pre-training.

Benefits

  • Top-tier compensation and equity.
  • Comprehensive medical, dental, vision, life, and disability insurance.
  • Fully paid parental leave, including adoptive and surrogate journeys, plus financial support for family planning.
  • Paid time off and relocation support.
  • Daily lunch and dinner, regular off-sites, and team celebrations.
Reflection

About Reflection

51-200 employees

Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. Our team previously built frontier LLMs at labs like DeepMind, OpenAI, and Anthropic. We believe AI should be built in the open, with transparent research and collaborative development. That means giving enterprises, governments, and sovereign entities true ownership and control of AI that performs at the highest level. Our mission: make intelligence open and accessible to all.