3 months ago
Base Salary
$180k - $360k/yr
Responsibilities
- Develop infrastructure and orchestration systems for deploying and managing large-scale distributed LLM inference.
- Build customer-facing features, inference libraries, and low-level infrastructure components across the stack.
- Create platform capabilities for routing, autoscaling, scheduling, observability, and runtime management.
- Improve the reliability, scalability, and usability of the inference stack.
- Collaborate with Model Performance engineers to make inference optimizations broadly available and easy to configure.
- Define best practices for testing, release automation, benchmarking, and operational excellence.
- Debug production systems spanning Kubernetes, distributed runtimes, networking, and GPU workloads.
- Make engineering tradeoffs among performance, reliability, operational simplicity, and developer experience.
- Own projects from architecture and implementation through deployment, monitoring, and iteration based on customer feedback.
Requirements
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, or a related field.
- Strong background in distributed systems, backend infrastructure, or platform engineering.
- Experience building and operating production systems where reliability, latency, and scale are first-class concerns.
- Strong focus on developer experience and how systems are used.
- Ability and willingness to learn new languages, frameworks, and systems.
- Ability to debug complex systems across multiple layers of the stack.
- Genuine interest in inference engineering; hands-on experience is not required.
- Excellent communication and collaboration skills.
- Experience with Kubernetes, including operators and custom resources, is a bonus.
- Prior work with Dynamo, vLLM, SGLang, TensorRT-LLM, or similar inference frameworks is a bonus.
- Experience with distributed scheduling, autoscaling, service orchestration, GPU workloads, observability tooling, CI/CD systems, release automation, or open-source infrastructure and ML systems is a bonus.
Benefits
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible PTO policy and company-wide Winter Break from Christmas Eve to New Year's Day.
- Paid parental leave.
- Fertility and family-building stipend through Carrot.
- Company-facilitated 401(k).
- Exposure to a variety of ML startups and learning and networking opportunities.
Tech Stack
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
