SpaceX

Sr. Site Reliability Engineer, AI Infrastructure (Starshield)

SpaceX
Apply
2 hours ago
Palo Alto, CA, USASenior
H1B sponsor

Base Salary

$165k - $265k/yr

Responsibilities

  • Manage GPU and CPU infrastructure deployments to Top Secret data centers.
  • Provide GPU-as-a-service support for external customers on bare-metal and virtualized platforms.
  • Design, validate, and productize AI cluster solutions at 100,000+ GPU scale.
  • Develop automation for deploying and managing on-premise Kubernetes and AI clusters and operating systems.
  • Deploy and manage databases, monitoring systems, and distributed storage.
  • Collaborate with AI engineers to build scalable, operable, and maintainable products.
  • Improve the full service lifecycle from design and deployment through operation and refinement.
  • Support monitoring and alerting systems to maintain high availability.
  • Identify reliability improvements and create innovative solutions for system availability.
  • Mentor junior engineers and lead the team toward technical excellence.

Requirements

  • Bachelor’s degree in computer science, information systems/IT, or engineering plus 5+ years of professional Linux experience, or 7+ years of software, DevOps, or site reliability engineering experience in lieu of a degree.
  • 5+ years of experience with Kubernetes and managing Linux operating systems.
  • Experience with Terraform, Ansible, or comparable infrastructure tools.
  • Experience with containerization technologies such as OCI containers and Kubernetes.
  • Scripting experience with Bash, Python, or similar languages.
  • Development experience in Python, C++, or Go.
  • Preferred qualifications include 5+ years of Python and Python-based development experience, Kubernetes cluster management, Linux boot and systems configuration knowledge, distributed databases and data modeling, large-scale server automation, TCP/IP networking, cloud virtualization, and NVIDIA GPU deployment stacks.
  • Preferred qualifications also include knowledge of Bazel, Makefiles, testing, continuous integration, build and deployment systems, continuous monitoring, and performance optimization.
  • Active Top Secret, Top Secret SCI, or DOE Level Q clearance is preferred.
  • Must be willing to work extended hours and weekends, travel domestically and globally when needed, and successfully obtain and maintain a Top Secret Security Clearance.
  • Applicants must satisfy applicable ITAR eligibility requirements or be eligible to obtain required U.S. Department of State authorizations.

Benefits

  • Annual base salary range of $165,000.00-$265,000.00, with potential stock or long-term cash awards, discretionary bonuses, and Employee Stock Purchase Plan participation.
  • Comprehensive medical, vision, and dental coverage; 401(k); disability and life insurance; paid parental leave; and other discounts and perks.
  • Approximately three weeks of paid vacation, 10 or more paid holidays, and paid sick leave under company policy.
  • Active clearance holders may receive a 10% differential, up to an additional $20,000 annually, after being briefed into a classified program.
  • The role may require extended hours, weekend work, and domestic or global travel.

Tech Stack

Categories

DevOpsSite Reliability
SpaceX

About SpaceX

10,000+ employees

SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.

Contact me