Cognition

Research Engineer, ML Infrastructure

Cognition
Apply
28 days ago

Responsibilities

  • Build and own distributed training infrastructure for large-scale jobs across GPU clusters, including launchers, checkpointing, recovery, fault tolerance, and monitoring.
  • Develop infrastructure for hundreds of thousands of concurrent coding-agent rollouts in VM sandboxes.
  • Profile and optimize training throughput, data loading, communication, memory utilization, compute efficiency, step time, and MFU.
  • Design and maintain experiment orchestration and research tooling for launching, tracking, and analyzing experiments.
  • Build reliable, high-throughput training and evaluation data pipelines with strong data quality, reproducibility, and efficiency.
  • Diagnose and resolve failures across GPUs, networking, numerics, and data, and build systems that recover quickly.
  • Implement and optimize data, tensor, pipeline, and sequence parallelism strategies for large-model training.
  • Anticipate future research infrastructure needs and scale systems ahead of demand.

Requirements

  • Demonstrated experience building and operating distributed training systems for large models, from cluster infrastructure through the training loop.
  • Strong systems engineering fundamentals in distributed systems, networking, storage, and full-stack hardware-software performance analysis.
  • Proficiency in Python and C++.
  • Experience with PyTorch or an equivalent deep learning framework at a systems level.
  • Hands-on experience with GPU performance profiling, memory optimization, compute efficiency, and diagnosing underperforming training runs.
  • Experience implementing or optimizing data, tensor, pipeline, or sequence parallelism for large-model training.
  • A track record of building tooling and abstractions that accelerate research workflows.
  • Strong debugging ability across complex, distributed, and nondeterministic systems.
  • Sufficient machine learning knowledge to collaborate substantively with researchers and understand training architectures and infrastructure needs.
  • Demonstrated capability is valued over formal credentials; a PhD is considered one possible signal rather than a stated requirement.

Benefits

  • Work on a small, highly selective team where research and product development move together and prototypes reach deployment quickly.
  • Own infrastructure operating across thousands of GPUs with substantial access to compute and systems.
  • Work in a fast-moving environment that rewards speed, autonomy, and technical depth with minimal process overhead.
  • Cognition is an equal opportunity employer and provides reasonable accommodations throughout the hiring process.

Categories

Data EngineeringML Engineering
Cognition

About Cognition

51-200 employees

Makers of Devin, the first AI software engineer. We are an applied AI lab building end-to-end software agents. We’re building collaborative AI teammates that enable engineers to focus on more interesting problems and empower engineering teams to strive for more ambitious goals.

Contact me