
Machine Learning Engineer - Inference
Together AI5 months ago
Base Salary
$160k - $230k/yr
Responsibilities
- Design and build production systems powering the Together AI inference engine for reliability and performance at scale.
- Develop and optimize runtime inference services for large-scale AI applications.
- Collaborate with researchers, engineers, product managers, and designers to deliver features and research capabilities.
- Conduct design and code reviews.
- Create services, tools, and developer documentation for the inference engine.
- Implement robust, fault-tolerant systems for data ingestion and processing.
Requirements
- At least 3 years of experience writing high-performance, well-tested, production-quality code.
- Proficiency with Python and PyTorch.
- Experience building high-performance libraries and tooling.
- Strong understanding of low-level operating system concepts, including multithreading, memory management, networking, storage, performance, and scale.
- Knowledge of AI inference systems such as TGI, vLLM, TensorRT-LLM, or Optimum is preferred.
- Knowledge of AI inference techniques such as speculative decoding is preferred.
- Knowledge of CUDA and Triton programming is preferred.
- Knowledge of Rust, Cython, and compilers is a nice-to-have.
Benefits
- Competitive compensation, startup equity, health insurance, and other benefits.
Categories
About Together AI
Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.