28 days ago
Responsibilities
- Build and own distributed training infrastructure for large-scale jobs across GPU clusters, including launchers, checkpointing, recovery, fault tolerance, and monitoring.
- Develop infrastructure for hundreds of thousands of concurrent coding-agent rollouts in VM sandboxes.
- Profile and optimize training throughput, data loading, communication, memory utilization, compute efficiency, step time, and MFU.
- Design and maintain experiment orchestration and research tooling for launching, tracking, and analyzing experiments.
- Build reliable, high-throughput training and evaluation data pipelines with strong data quality, reproducibility, and efficiency.
- Diagnose and resolve failures across GPUs, networking, numerics, and data, and build systems that recover quickly.
- Implement and optimize data, tensor, pipeline, and sequence parallelism strategies for large-model training.
- Anticipate future research infrastructure needs and scale systems ahead of demand.
Requirements
- Demonstrated experience building and operating distributed training systems for large models, from cluster infrastructure through the training loop.
- Strong systems engineering fundamentals in distributed systems, networking, storage, and full-stack hardware-software performance analysis.
- Proficiency in Python and C++.
- Experience with PyTorch or an equivalent deep learning framework at a systems level.
- Hands-on experience with GPU performance profiling, memory optimization, compute efficiency, and diagnosing underperforming training runs.
- Experience implementing or optimizing data, tensor, pipeline, or sequence parallelism for large-model training.
- A track record of building tooling and abstractions that accelerate research workflows.
- Strong debugging ability across complex, distributed, and nondeterministic systems.
- Sufficient machine learning knowledge to collaborate substantively with researchers and understand training architectures and infrastructure needs.
- Demonstrated capability is valued over formal credentials; a PhD is considered one possible signal rather than a stated requirement.
Benefits
- Work on a small, highly selective team where research and product development move together and prototypes reach deployment quickly.
- Own infrastructure operating across thousands of GPUs with substantial access to compute and systems.
- Work in a fast-moving environment that rewards speed, autonomy, and technical depth with minimal process overhead.
- Cognition is an equal opportunity employer and provides reasonable accommodations throughout the hiring process.
Categories
Data EngineeringML Engineering
About Cognition
Makers of Devin, the first AI software engineer. We are an applied AI lab building end-to-end software agents. We’re building collaborative AI teammates that enable engineers to focus on more interesting problems and empower engineering teams to strive for more ambitious goals.
