6 months ago
Base Salary
$225k - $550k/yr
Responsibilities
- Design and scale high-performance inference serving systems
- Optimize KV-cache management, batching strategies, and scheduling
- Improve throughput and latency for long-context workloads
- Build and maintain distributed RL and post-training infrastructure
- Improve reliability of rollout, evaluation, and reward pipelines
- Automate fault detection and recovery for serving and RL systems
- Profile and eliminate performance bottlenecks across GPU, networking, and storage layers
- Collaborate with Kernels and Research to align execution systems with model architecture
Requirements
- Strong software engineering and distributed systems fundamentals
- Experience building or operating large-scale inference or training systems
- Deep understanding of GPU execution constraints and memory trade-offs
- Experience debugging performance issues in production ML systems
- Ability to reason about latency, throughput, and cost trade-offs
- Track record of owning critical production infrastructure
Benefits
- Equity in addition to salary
- 401(k) plan with 6% salary matching
- Health, dental, and vision insurance for employees and dependents
- Unlimited paid time off
- Visa sponsorship and relocation stipend to San Francisco, if possible
- Small, fast-paced, highly focused team
Categories
About Magic
Magic is working on frontier-scale code models to build a coworker, not just a copilot. Come join us: http://magic.dev
