6 hours ago
Base Salary
$125k - $200k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret datacenters.
- Provide GPU-as-a-service support for external customers using bare-metal hardware and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000+ GPU scale.
- Develop automation for on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to build scalable, operable, and maintainable products.
- Improve the full service lifecycle from design and deployment through operation and refinement.
- Implement monitoring and alerting and improve system availability.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or engineering plus 1+ years of professional site reliability engineering or DevOps experience, or 3+ years of such experience in lieu of a degree.
- At least 1 year of professional experience with Linux operating systems.
- Experience with Terraform, Ansible, or similar infrastructure tools.
- Experience with containerization technologies such as OCI containers and Kubernetes.
- Scripting experience with Bash, Python, or similar languages.
- Development experience in Python, C++, or Go.
- Preferred experience includes Python-based development frameworks, managing Kubernetes clusters, Linux systems configuration, testing and continuous monitoring, Bazel, Makefiles, distributed databases, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
- An active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred.
- Must be able to obtain and maintain a Top Secret Security Clearance.
- Must be willing to work extended hours and weekends and travel domestically and globally when needed.
Benefits
- Base salary range is $125,000-$200,000 depending on level, knowledge, skills, education, and experience.
- Potential long-term incentives, long-term cash awards, discretionary bonuses, and Employee Stock Purchase Plan participation.
- Medical, vision, dental, 401(k), disability, life insurance, paid parental leave, discounts, and other perks.
- Approximately three weeks of paid vacation and 10 or more paid holidays annually.
- Washington employees accrue paid sick time in compliance with applicable law.
- Company shuttles operate Monday through Friday from select Seattle locations to the SpaceX Redmond office.
- Employees with an active clearance may receive a 10% differential, up to an additional $20,000 annually, after being briefed into a classified program.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
