7 hours ago
Base Salary
$165k - $265k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000+ GPU scale.
- Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to build scalable, operable, and maintainable products.
- Improve services throughout their lifecycle, from design and deployment through operation and refinement.
- Maintain monitoring and alerting systems and improve high availability.
- Identify infrastructure improvements and create solutions that increase system availability.
- Mentor junior engineers and lead the team toward technical excellence.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus 5+ years of professional Linux experience, or 7+ years of professional experience in software, DevOps, or site reliability engineering in lieu of a degree.
- 5+ years of experience with Kubernetes and managing Linux operating systems.
- Experience with Terraform, Ansible, or comparable infrastructure tools and with containerization technologies such as OCI containers and Kubernetes.
- Scripting experience with Bash, Python, or similar languages, plus development experience in Python, C++, or Go.
- Preferred qualifications include 5+ years of Python and Python-based development, Kubernetes cluster management, Linux boot and systems configuration knowledge, testing and continuous integration/build/deployment/monitoring knowledge, Bazel or Makefiles, performance optimization, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
- Strong communication skills and an active Top Secret, Top Secret SCI, or DOE Level Q clearance are preferred.
- Must be willing to work extended hours and weekends, travel domestically and globally as needed, and successfully obtain and maintain a Top Secret security clearance.
Benefits
- Base salary range is $165,000.00-$265,000.00 annually, with potential stock or long-term cash awards, discretionary bonuses, and Employee Stock Purchase Plan participation.
- Benefits include medical, vision, and dental coverage, a 401(k), disability and life insurance, paid parental leave, discounts, approximately three weeks of paid vacation, at least 10 paid holidays, and paid sick leave.
- Employees with an active clearance may receive a 10% differential, up to an additional $20,000 annually, after being briefed into a classified program.
- The role may require extended hours, weekends, and domestic or global travel; obtaining and maintaining a Top Secret clearance is a condition of employment.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
