Magic

Member of Technical Staff, Pre-training Systems

Magic
Apply
6 months ago

Base Salary

$225k - $550k/yr

Responsibilities

  • Scale distributed training across large GPU clusters using data, tensor, and pipeline parallelism.
  • Optimize communication patterns and gradient synchronization.
  • Improve checkpointing, fault tolerance, and job recovery systems.
  • Profile and eliminate performance bottlenecks across compute, networking, and storage.
  • Improve experiment reproducibility and orchestration workflows.
  • Increase hardware utilization and training throughput.
  • Collaborate with Kernels and Research to align model architecture with systems realities.

Requirements

  • Strong software engineering and distributed systems fundamentals.
  • Experience training large models in multi-node GPU environments.
  • Deep understanding of parallelism strategies and performance trade-offs.
  • Experience debugging cross-layer issues in production machine learning systems.
  • Strong ownership mindset and ability to operate critical infrastructure.
  • Track record of improving performance or reliability of large-scale systems.

Benefits

  • Annual salary range of $225K-$550K plus significant equity compensation.
  • 401(k) plan with 6% salary matching.
  • Health, dental, and vision insurance for employees and dependents.
  • Unlimited paid time off.
  • Visa sponsorship and relocation stipend to San Francisco, if possible.
  • Small, fast-paced, highly focused team environment.
Magic

About Magic

51-200 employees

Magic is working on frontier-scale code models to build a coworker, not just a copilot. Come join us: http://magic.dev