over 2 years ago
Base Salary
$200k - $550k/yr
Responsibilities
- Design and operate large-scale GPU clusters for training and inference.
- Build and maintain Terraform infrastructure across cloud and hybrid environments.
- Deploy, operate, and optimize Kubernetes clusters used to schedule and manage AI workloads.
- Develop modular, scalable infrastructure-as-code patterns for compute, networking, and storage provisioning.
- Improve deployment reproducibility, environment consistency, and operational safety.
- Optimize networking and storage systems for high-throughput AI workloads.
- Automate fault detection and recovery across distributed clusters.
- Debug cross-layer issues involving hardware, drivers, networking, storage, operating systems, and cloud environments.
- Improve observability, monitoring, and reliability of core platform systems.
Requirements
- Strong systems engineering fundamentals.
- Deep hands-on experience with Terraform, including module design, state management, environment isolation, and large-scale deployments.
- Experience operating production GPU infrastructure or high-performance distributed systems.
- Strong understanding of networking and storage systems.
- Experience with major cloud platforms such as GCP, AWS, Azure, or OCI.
- Track record of owning production-critical infrastructure end-to-end.
Benefits
- Annual salary range of $200K-$550K depending on experience.
- Equity is included as a significant part of total compensation.
- 401(k) plan with 6% salary matching.
- Health, dental, and vision insurance for the employee and dependents.
- Unlimited paid time off.
- Visa sponsorship and relocation stipend to San Francisco, if possible.
- Small, fast-paced, highly focused team.
Categories
About Magic
Magic is working on frontier-scale code models to build a coworker, not just a copilot. Come join us: http://magic.dev
