almost 2 years ago
Base Salary
$180k - $250k/yr
Responsibilities
- Design and build low-latency, scalable, and reliable model inference and serving systems for foundation models using Transformers, SSMs, and hybrid models.
- Collaborate with research and product engineering teams to serve Cartesia’s products quickly, cost-effectively, and reliably.
- Build robust inference infrastructure and monitoring for products.
- Help shape products and apply cutting-edge AI across devices and applications.
Requirements
- Strong engineering skills and the ability to navigate complex codebases while writing clean, maintainable code.
- Experience building large-scale distributed systems with demanding performance, reliability, and observability requirements.
- Technical leadership and the ability to deliver zero-to-one results amid ambiguity.
- Background in or experience with inference pipelines for machine-learning and generative models.
- Experience implementing state-of-the-art machine-learning models and applying research to practical problems.
- Experience with vLLM, SGLang, continuous batching, or other inference frameworks is preferred.
- Experience with CUDA, Triton, or similar technologies is preferred.
Benefits
- In-person work from offices in San Francisco, London, or Bangalore.
- Visa sponsorship support is assessed case by case.
- Competitive base salary with an attractive equity package.
- Monthly commuter allowance.
- Flexible PTO.
- Daily lunch, dinner, and snacks.
Categories
About Cartesia
Cartesia builds real-time audio AI models and APIs for developers creating interactive voice agents and applications. Its products include Sonic for low-latency speech recognition, text-to-speech, and speech-to-speech, and Ink as companion tooling for interactive agents. Founded in 2023, the company is privately held and offers its technology as an AI/API platform for software teams.
