SpaceX

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Apply
4 hours ago
Washington, DC, USASenior
H1B sponsor

Base Salary

$165k - $265k/yr

Responsibilities

  • Manage GPU and CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support for external customers across bare-metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions at 100,000-plus GPU scale.
  • Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
  • Deploy and manage databases, monitoring systems, distributed storage, and other core infrastructure.
  • Collaborate with AI engineers to build scalable, operable, and maintainable products.
  • Manage the full service lifecycle from design and deployment through operation and refinement.
  • Improve monitoring, alerting, system availability, and performance.
  • Mentor junior engineers and lead the team toward technical excellence.

Requirements

  • Bachelor’s degree in computer science, information systems/IT, or engineering plus 5 or more years of professional Linux experience, or 7 or more years of professional software, DevOps, or site reliability engineering experience in lieu of a degree.
  • At least 5 years of Kubernetes experience and at least 5 years managing Linux operating systems.
  • Experience with Terraform, Ansible, or comparable infrastructure tools and with containerization technologies such as OCI containers and Kubernetes.
  • Experience scripting in Bash, Python, or similar languages and developing software in Python, C++, or Go.
  • Preferred: 5 or more years of Python and Python-based development frameworks, Kubernetes cluster management, Linux boot and systems configuration, and distributed databases and data modeling.
  • Preferred: experience with testing, continuous integration, build and deployment technologies, Bazel, Makefiles, automated management of thousands of servers, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks including Blackwell and Rubin.
  • Preferred: active Top Secret, Top Secret SCI, or DOE Level Q clearance; the role requires obtaining and maintaining a Top Secret clearance.
  • Must be willing to work extended hours and weekends, travel domestically and globally as needed, and communicate effectively with customers, peers, and management.

Benefits

  • Base salary range is $165,000 to $265,000 annually, with potential long-term incentives, bonuses, and employee stock purchase opportunities.
  • Benefits include medical, vision, dental, 401(k), short- and long-term disability insurance, life insurance, paid parental leave, discounts, three weeks of accrued paid vacation, paid holidays, and paid sick leave.
  • Employees with an active clearance may receive a 10% differential up to an additional $20,000 annually after being briefed into a classified program.
  • The position may require extended hours, weekend work, and domestic or global travel.
  • Employment is subject to U.S. ITAR eligibility requirements and obtaining a Top Secret security clearance.

Tech Stack

Categories

DevOpsSite Reliability
SpaceX

About SpaceX

10,000+ employees

SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.

Contact me