Together AI

Senior Backend Engineer, Inference Platform

Together AI
Apply
4 months ago

Base Salary

$160k - $250k/yr

Responsibilities

  • Build and optimize global and local request routing and low-latency load balancing across data centers and model engine pods.
  • Develop autoscaling systems that dynamically allocate resources across dozens of data centers while meeting strict SLOs.
  • Design multi-tenant traffic shaping, rate limiting, regulation, and resource allocation systems.
  • Optimize latency, throughput, prefix caching, and system-level performance for diverse workloads.
  • Collaborate with ML researchers to bring new model architectures into production at scale.
  • Profile systems, identify bottlenecks, and implement performance optimizations.
  • Collaborate with and contribute to the open source inference community.

Requirements

  • At least 5 years of experience building large-scale, fault-tolerant distributed systems and API microservices.
  • Strong background designing and improving the efficiency, scalability, and stability of complex systems.
  • Excellent understanding of multithreading, memory management, networking, and storage performance.
  • Expert-level programming in one or more of Rust, Go, Python, or TypeScript.
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
  • Knowledge of modern LLMs and generative models and how they are served in production is a plus.
  • Experience with the open source inference ecosystem, especially SGLang, vLLM, or NVIDIA Dynamo, is valuable.
  • Experience with Kubernetes or container orchestration is a strong plus.
  • Familiarity with CUDA, Triton, NCCL, InfiniBand, NVLink, or MPI is a plus.

Benefits

  • Competitive compensation, equity, health insurance, and other benefits.
  • Full-time position.
Together AI

About Together AI

201-500 employees

Together AI builds an AI-native cloud platform for developers, offering high-performance inference, fine-tuning/model shaping, and large-scale pre-training on on-demand GPU clusters with APIs and managed services. It emphasizes open-source models that teams can run and adapt, and also provides infrastructure for decentralized and scalable workloads. Founded in 2022 and headquartered in San Francisco, it is privately held and reports notable customers including Cursor, ElevenLabs, Salesforce, and Zoom.

Contact me