
Member of Technical Staff - Compute Platform
Reflection7 months ago
Responsibilities
- Build and maintain tools for automatic remediation, topology-aware scheduling, capacity planning, and rapid hardware debugging
- Design and iterate on the cluster management stack for workloads across large multi-GPU fleets
- Implement cluster-wide monitoring focused on durability and active performance benchmarking
- Prepare infrastructure for next-generation GPU deployments and increasingly larger cluster sizes
- Help own multi-cloud storage, petabyte-scale data replication, and GPU-to-GPU network performance
- Collaborate with training teams to co-design fault tolerance, node health checks, and remediation strategies
Requirements
- Systems-level engineering experience focused on cluster-wide behavior and maintenance
- Strong coding ability with demonstrated systems or GPU infrastructure experience
- Deep GPU hardware knowledge beyond standard Kubernetes, including familiarity with NCCL
- Alignment with a Kubernetes-first architecture
- Cloud storage expertise, including managing high-performance data products such as VAST across multiple data centers and handling datasets and checkpointing at scale
Benefits
- Comprehensive medical, dental, vision, life, and disability insurance
- Fully paid parental leave for all new parents, including adoptive and surrogate journeys
- Financial support for family planning
- Paid time off
- Relocation support
- Daily provided lunch and dinner
- Regular off-sites and team celebrations
Tech Stack
Categories
DevOpsSite Reliability
About Reflection
Reflection is a New York–based, privately held research lab developing open foundational AI models and agentic coding tools for developers, enterprises, and public-sector users. The team includes former researchers from DeepMind, OpenAI, and Anthropic, and their work focuses on transparent, customizable systems that organizations can deploy with ownership and control.