2 months ago
Base Salary
$135k - $210k/yr
Responsibilities
- Develop reliable, high-throughput inference systems serving AI models internally across SpaceX.
- Architect and implement scalable distributed model-serving infrastructure, including load balancing, autoscaling, batch scheduling, global KV caching, and continuous batching.
- Optimize inference latency and throughput through GPU kernel work, quantization, speculative decoding, and other acceleration techniques.
- Build high-concurrency serving systems with strong availability, low tail latency, and observability.
- Own inference-platform components including request routing, SDKs, rate limiting, and scaling.
- Benchmark, fine-tune, and accelerate inference engines including SGLang, vLLM, and TensorRT-LLM.
- Develop tools for tracing, replaying, and resolving issues across orchestration and GPU-kernel layers.
- Create CI/CD infrastructure for endpoint deployment, image publishing, and inference-engine updates.
- Collaborate with SpaceX AI teams to integrate inference capabilities into broader systems and workflows.
Requirements
- Bachelor’s degree in computer science, engineering, mathematics, or a scientific discipline, or 2+ years of professional software-building experience in lieu of a degree.
- Experience designing, implementing, and maintaining reliable, horizontally scalable distributed systems.
- At least 1 year of professional full-stack or backend development experience with production systems.
- At least 1 year of experience with Rust or C++.
- Preferred experience with LLM inference engines and serving frameworks such as SGLang, vLLM, Triton, and TensorRT-LLM.
- Preferred experience with GPU kernels, code generation, batching, caching, parallelism, quantization, speculative decoding, and large-scale high-concurrency serving systems.
- Knowledge of service observability and reliability practices, profiling, and application performance improvement.
- Experience with PostgreSQL, ClickHouse, or MongoDB.
- Experience with agent SDKs, agent orchestration frameworks, Docker, Kubernetes, and containerized applications.
- Expertise in gRPC, including unary, response streaming, bidirectional streaming, and REST mapping; programming experience in Python, Go, or similar languages.
- Applicants must meet applicable ITAR eligibility requirements or be eligible to obtain required authorizations.
Benefits
- Base salary is offered at Level 1 or Level 2, with additional eligibility for company stock or long-term cash awards, discretionary bonuses, and an employee stock purchase plan.
- Comprehensive medical, vision, and dental coverage is provided, along with a 401(k), disability insurance, life insurance, paid parental leave, discounts, and other perks.
- Employees may accrue 3 weeks of paid vacation and receive 10 or more paid holidays annually, plus paid sick leave under company policy.
- The role requires onsite work in Palo Alto; remote and hybrid work are not considered.
- The role may require extended hours or weekend work depending on launch cadence and platform demands.
Tech Stack
Categories
About SpaceX
SpaceX designs, manufactures, and launches orbital rockets and spacecraft, and operates Starlink, a global satellite internet network for consumers, businesses, and governments. Its revenue comes from commercial and government launch services (Falcon 9/Falcon Heavy, Dragon cargo and crew to the ISS) and subscription broadband with Starlink hardware and service. Founded in 2002 and headquartered in Hawthorne, California, the privately held company serves NASA and commercial satellite operators, building most systems in-house.
