7 hours ago
Base Salary
$165k - $270k/yr
Responsibilities
- Manage GPU and CPU infrastructure deployments to Top Secret data centers.
- Support GPU-as-a-service for external customers on bare-metal and virtualized platforms.
- Design, validate, and productize AI cluster solutions at 100,000+ GPU scale.
- Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
- Deploy and manage databases, monitoring systems, and distributed storage.
- Collaborate with AI engineers to create scalable, operable, and maintainable products.
- Improve services throughout their lifecycle, from design and deployment through operation and refinement.
- Maintain monitoring and alerting systems supporting high availability.
- Identify infrastructure improvements and create solutions that increase system availability.
- Mentor and train junior engineers and lead the team toward technical excellence.
Requirements
- Bachelor’s degree in computer science, information systems/IT, or an engineering discipline plus 5+ years of professional Linux experience, or 7+ years of professional software, DevOps, or site reliability engineering experience in lieu of a degree.
- 5+ years of experience with Kubernetes and Linux operating systems.
- Experience with Terraform, Ansible, or comparable infrastructure tools.
- Experience with containerization technologies such as OCI containers and Kubernetes.
- Scripting experience with Bash, Python, or similar languages.
- Development experience in Python, C++, or Go.
- Preferred: 5+ years of Python and Python-based development framework experience.
- Preferred: experience managing Kubernetes clusters, Linux boot and system configuration, distributed databases and data modeling, large-scale server management, cloud virtualization, and NVIDIA GPU deployment stacks.
- Preferred: knowledge of testing, continuous integration, build, deployment, continuous monitoring, Bazel, Makefiles, performance optimization, and TCP/IP networking.
- Preferred: active Top Secret, Top Secret SCI, or DOE Level Q clearance.
- Must be willing to obtain and maintain a Top Secret security clearance.
- Must be willing to work extended hours and weekends as needed and travel domestically and globally when required.
- Applicants must meet applicable ITAR eligibility requirements or be eligible to obtain required U.S. Department of State authorizations.
Benefits
- Base salary range is $165,000.00-$270,000.00, with possible long-term incentives, long-term cash awards, discretionary bonuses, and employee stock purchase opportunities.
- Comprehensive medical, vision, and dental coverage; 401(k); disability and life insurance; paid parental leave; discounts and other perks.
- Approximately three weeks of paid vacation and 10 or more paid holidays annually.
- Washington State employees accrue paid sick time in accordance with state and federal law.
- Company shuttles are available Monday through Friday from select Seattle locations to the SpaceX Redmond office.
- Employees with an active clearance receive a 10% differential, up to an additional $20,000 annually, once briefed into a classified program.
- The role may require extended hours, weekends, and domestic or global travel.
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
