4 months ago
Base Salary
$165k - $330k/yr
Responsibilities
- Own and lead Voice AI product areas end-to-end from architecture and system design through implementation, rollout, and long-term production operations
- Design, build, and operate real-time, large-scale, high-performance model-serving systems for STT, TTS, and voice-agent workloads
- Optimize end-to-end and tail latency, throughput, and GPU efficiency through profiling, runtime tuning, and server-level optimizations
- Build streaming infrastructure that orchestrates multi-model voice-agent components and meets customer SLOs
- Create training and inference iteration loops for voice-model customization, evaluation, rollout, and experimentation
- Collaborate with sister engineering teams and Forward Deployed Engineers to solve full-stack problems and coordinate delivery
- Mentor teammates through code reviews, design documents, and technical leadership
Requirements
- Bachelor's degree or higher in computer science or a related field
- Proven track record owning production-grade, real-time, large-scale systems where p99 tail latency matters
- Proficient coding abilities in one or more popular programming or scripting languages; Python proficiency is a plus
- Good product judgment, particularly for developer-oriented tools
- Interest in ML/AI infrastructure and willingness to learn
- Strong collaboration and communication skills
- Comfort using AI coding assistants such as Claude Code, Codex, or Cursor
Benefits
- 100% medical, dental, and vision insurance coverage for employees and dependents
- Flexible PTO and company-wide Winter Break from Christmas Eve through New Year's Day
- Paid parental leave
- Fertility and family-building stipend through Carrot
- Company-facilitated 401(k)
- Meaningful equity and competitive compensation
- Exposure to a variety of ML startups and related learning and networking opportunities
Tech Stack
Categories
About Baseten
Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.
