6 months ago
Berlin, GermanySenior / Staff+
Responsibilities
- Optimize model inference and serving for latency, throughput, and reliability.
- Implement model acceleration techniques including quantization, distillation, caching, continuous batching, paged attention, and speculative decoding.
- Build and scale distributed inference systems across multiple GPUs and nodes.
- Profile high-performance systems and improve utilization of NVIDIA GPUs.
- Containerize research models, optimize their serving, and operate them reliably in production.
- Clarify ambiguous problems through benchmarks and prototypes and take full-cycle ownership of shipped systems.
Requirements
- Deep understanding of modern serving frameworks and inference optimization techniques such as vLLM or TRT-LLM.
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding.
- Proficiency in C++, CUDA, Rust, or highly optimized Python, including profiling and optimizing performance on NVIDIA GPUs.
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and handling thousands of concurrent connections.
- Public technical work such as substantial systems programming projects, open-source contributions to inference engines, or deep technical write-ups is useful.
- Ability to take a model from research through containerization, serving optimization, and reliable production operation.
- PhD in computer science, physics, mathematics, or equivalent practical experience building backend or machine-learning systems.
- Professional fluency in written and spoken English.
Benefits
- Full U.S. visa and relocation support may be available for candidates interested in relocating to the San Francisco Bay Area, subject to business needs and work authorization requirements.
Tech Stack
Categories
About Inworld
Inworld builds real-time voice AI and agent infrastructure for developers, including text-to-speech, speech-to-speech, speech recognition, and LLM routing delivered via APIs and SDKs. Customers use it to create interactive characters and agentic experiences in games, apps, and virtual worlds; revenue comes from usage-based and enterprise licensing. Founded in 2021 and headquartered in Mountain View, California, the company is privately held and works with game studios and large technology firms.
