
Member of Technical Staff - Compute Platform
Reflection5 months ago
Responsibilities
- Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging
- Design and iterate on the cluster management stack for workloads across large multi-GPU fleets
- Implement cluster-wide monitoring focused on durability and active performance benchmarking
- Prepare infrastructure for next-generation GPU deployments and increasingly larger cluster sizes
- Help own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance
- Collaborate with training teams to co-design fault tolerance, node health checks, and remediation strategies
Requirements
- Systems-level engineering experience focused on cluster-wide behavior and maintenance
- Strong coding ability with demonstrated systems or GPU infrastructure experience
- Deep GPU hardware knowledge beyond standard Kubernetes, including familiarity with NCCL
- Alignment with a Kubernetes-first architecture
- Cloud storage expertise, including managing high-performance data products such as VAST across multiple data centers and handling datasets and checkpointing at scale
Benefits
- Comprehensive medical, dental, vision, life, and disability insurance
- Fully paid parental leave for all new parents, including adoptive and surrogate journeys
- Financial support for family planning
- Paid time off
- Relocation support
- Daily provided lunch and dinner
- Regular off-sites and team celebrations
Tech Stack
Categories
DevOpsSite Reliability
About Reflection
Reflection is a research lab making intelligence open and accessible for everyone to use, customize, and build on. Our team previously built frontier LLMs at labs like DeepMind, OpenAI, and Anthropic. We believe AI should be built in the open, with transparent research and collaborative development. That means giving enterprises, governments, and sovereign entities true ownership and control of AI that performs at the highest level. Our mission: make intelligence open and accessible to all.