7 months ago
Base Salary
$225k - $550k/yr
Responsibilities
- Scale distributed training across large GPU clusters using data, tensor, and pipeline parallelism.
- Optimize communication patterns and gradient synchronization.
- Improve checkpointing, fault tolerance, and job recovery systems.
- Profile and eliminate performance bottlenecks across compute, networking, and storage.
- Improve experiment reproducibility and orchestration workflows.
- Increase hardware utilization and training throughput.
- Collaborate with Kernels and Research to align model architecture with systems realities.
Requirements
- Strong software engineering and distributed systems fundamentals.
- Experience training large models in multi-node GPU environments.
- Deep understanding of parallelism strategies and performance trade-offs.
- Experience debugging cross-layer issues in production machine learning systems.
- Strong ownership mindset and ability to operate critical infrastructure.
- Track record of improving performance or reliability of large-scale systems.
Benefits
- Annual salary range of $225K-$550K plus significant equity compensation.
- 401(k) plan with 6% salary matching.
- Health, dental, and vision insurance for employees and dependents.
- Unlimited paid time off.
- Visa sponsorship and relocation stipend to San Francisco, if possible.
- Small, fast-paced, highly focused team environment.
Categories
About Magic
Magic builds frontier-scale code models designed to act as an AI coworker for software developers and engineering teams, automating code generation and research tasks. Its products center on developer-facing models and tooling that integrate into software workflows for teams seeking higher velocity and reliability. Founded in 2022 and headquartered in San Francisco, the privately held company focuses on AI-driven developer tools spanning information technology and machine learning.
