15 hours ago
Base Salary
$125k - $195k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Provide GPU-as-a-service support for external customers using bare-metal hardware and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
- Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to build scalable, operable, and maintainable products.
- Improve services across their lifecycle, from design and deployment through operation and refinement.
- Maintain monitoring and alerting systems and develop solutions that improve system availability.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus at least one year of professional site reliability engineering or DevOps experience, or at least three years of such experience in lieu of a degree.
- At least one year of professional experience with Linux operating systems.
- Experience with Terraform, Ansible, or comparable infrastructure tools.
- Experience with containerization technologies such as OCI containers and Kubernetes.
- Scripting experience in Bash, Python, or similar languages.
- Development experience in Python, C++, or Go.
- Preferred experience includes Python-based development frameworks, managing Kubernetes clusters, Linux boot and system configuration, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
- Preferred knowledge includes testing, continuous integration, build and deployment technologies, continuous monitoring, Bazel, Makefiles, and performance optimization.
- An active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred; the role requires successfully obtaining and maintaining a Top Secret clearance.
- Must be willing to work extended hours and weekends and travel domestically and globally as needed.
Benefits
- Base salary is listed in two levels ranging from $125,000 to $195,000 annually, with a 10% clearance differential up to an additional $20,000 annually; compensation figures are excluded from benefit details.
- Eligible employees may receive company stock or long-term cash awards, discretionary bonuses, and access to an Employee Stock Purchase Plan.
- Benefits include medical, vision, dental, 401(k), short- and long-term disability, life insurance, paid parental leave, discounts, three weeks of paid vacation, paid holidays, and paid sick leave.
- The role requires obtaining and maintaining a Top Secret security clearance and may require extended hours, weekend work, and future domestic or global travel.
- ITAR eligibility requires U.S. citizenship or national status, lawful permanent residence, refugee or asylee status, or eligibility to obtain required U.S. Department of State authorizations.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
