1 day ago
Responsibilities
- Optimize research models for production using sparsification, distillation, quantization, pruning, and mixed precision.
- Own model optimization lifecycles by defining metrics, running experiments, and benchmarking latency, cost, and quality trade-offs.
- Partner with researchers and engineers to turn new ideas into deployable systems.
- Analyze inference performance and GPU/accelerator efficiency for large multimodal models.
- Read ML papers, reproduce results, and adapt research ideas to practical systems.
Requirements
- Strong experience in deep learning with PyTorch.
- Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision.
- Understanding of efficient architectures such as low-rank adapters.
- Strong understanding of inference performance and GPU/accelerator fundamentals.
- Strong Python coding skills and reliable research engineering practices.
- Experience working with large models and datasets in cloud environments.
- Clear communication and collaboration skills.
- Preferred experience with diffusion models, video/audio generative models, or large language models.
- Preferred experience with real-time or streaming systems such as low-latency APIs, WebRTC, or streaming TTS/video.
- Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA.
- Experience writing custom Triton/CUDA kernels or performing low-level performance tuning.
- Experience with experiment tracking, benchmarking, and profiling at scale.
- Prior research engineering or applied science experience.
Benefits
- Flexible work schedules.
- Unlimited PTO.
- Competitive healthcare and gear stipends.
- Collaborative environment focused on learning and impact.
- Preferably hybrid in San Francisco, with remote candidates also considered.
- Relocation support is offered.
Categories
About Tavus
Tavus builds generative AI technology and APIs that let developers and enterprises create real-time, face-to-face conversational agents and digital-twin video experiences. Its platform exposes speech, vision, and video synthesis models for embedding into apps for customer engagement, healthcare assistants, training, and other interactive use cases, sold as usage-based developer services. Founded in 2020 and headquartered in San Francisco, the company is privately held.
