8 hours ago
Base Salary
$165k - $265k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Provide support for GPU-as-a-service platforms on bare-metal hardware and virtualized environments.
- Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
- Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, distributed storage, and other core infrastructure.
- Collaborate with AI engineers to build scalable, operable, and maintainable products.
- Manage the full service lifecycle from design and deployment through operation and refinement.
- Improve monitoring and alerting to support high availability.
- Identify infrastructure improvements and develop solutions that increase system availability.
- Mentor and train junior engineers and lead the team toward technical excellence.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or engineering plus at least five years of professional Linux experience, or at least seven years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree.
- At least five years of Kubernetes experience and at least five years of Linux operating-system management experience.
- Experience with Terraform, Ansible, or comparable infrastructure tools.
- Experience with OCI containers, Kubernetes, or other containerization technologies.
- Scripting experience with Bash, Python, or similar languages and development experience with Python, C++, or Go.
- Preferred qualifications include five or more years of Python and Python-based development, Kubernetes cluster management, Linux boot and systems configuration, testing and continuous integration, build and deployment technologies such as Bazel and Makefiles, performance optimization, distributed databases, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks including Blackwell and Rubin.
- Preferred experience includes automatically managing thousands of servers and active Top Secret, Top Secret SCI, or DOE Level Q clearance.
- Must be willing to work extended hours and weekends and travel domestically and globally as needed.
- Must successfully obtain and maintain a Top Secret security clearance as a condition of employment.
Benefits
- Base salary range is $165,000 to $265,000 annually, with potential long-term incentives, discretionary bonuses, and employee stock purchase plan participation.
- Benefits include medical, vision, dental, 401(k), disability and life insurance, paid parental leave, discounts, approximately three weeks of paid vacation, at least 10 paid holidays, and paid sick leave.
- Employees with an active clearance may receive a 10% differential up to an additional $20,000 annually after being briefed into a classified program.
- The role may require extended hours, weekend work, and future domestic and global travel.
- Employment is subject to U.S. ITAR eligibility requirements and required security-clearance maintenance.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
