4 hours ago
Base Salary
$165k - $265k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000+ GPU scale.
- Develop automation for on-premise Kubernetes/AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to create scalable, operable, and maintainable products.
- Improve service lifecycles from design and deployment through operation and refinement.
- Maintain monitoring and alerting systems that support high availability.
- Identify reliability improvements and develop innovative solutions for system availability.
- Mentor junior engineers and lead the team toward technical excellence.
- Work extended hours and weekends and travel domestically or globally when required.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or engineering plus 5+ years of professional Linux experience, or 7+ years of professional software, DevOps, or site reliability engineering experience in lieu of a degree.
- 5+ years of Kubernetes experience and 5+ years managing Linux operating systems.
- Experience with Terraform, Ansible, or comparable infrastructure tools.
- Experience with OCI containers, Kubernetes, or other containerization technologies.
- Scripting experience with Bash, Python, or similar languages and development experience with Python, C++, or Go.
- Preferred: 5+ years of Python and Python-based development framework experience.
- Preferred: Kubernetes cluster management, Linux boot and system configuration, testing and continuous integration, build and deployment technologies, and continuous monitoring knowledge.
- Preferred: experience with Bazel, Makefiles, distributed databases, data modeling, large-scale server management, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks including Blackwell and Rubin.
- Preferred: active Top Secret, Top Secret SCI, or DOE Level Q clearance.
- Must successfully obtain and maintain a Top Secret security clearance as a condition of employment.
- Must meet applicable ITAR eligibility requirements or be eligible to obtain required U.S. Department of State authorizations.
- Strong communication skills and the ability to communicate with customers, peers, and management.
Benefits
- Base salary range is $165,000.00–$265,000.00 for Level 3, with potential stock or long-term cash awards, discretionary bonuses, and an Employee Stock Purchase Plan.
- Benefits include medical, vision, and dental coverage; a 401(k); disability and life insurance; paid parental leave; discounts and other perks.
- Employees may accrue three weeks of paid vacation and receive 10 or more paid holidays annually, plus paid sick leave under company policy.
- Employees with an active clearance may receive a 10% differential, up to an additional $20,000 annually, after being briefed into a classified program.
- The role may require extended hours, weekends, and future domestic or global travel.
- The position requires obtaining and maintaining a Top Secret security clearance.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
