6 months ago
Base Salary
$270k - $500k/yr
Responsibilities
- Own the full cycle of machine learning model serving from research handoff to reliable production operation.
- Optimize inference latency and throughput using model acceleration and serving techniques.
- Profile and tune high-performance systems running on NVIDIA GPUs.
- Design and operate distributed, multi-GPU and multi-node inference systems.
- Build scaling, load-balancing, and reliability solutions for thousands of concurrent connections.
- Containerize models and ensure their serving infrastructure is stable in production.
- Develop benchmarks and prototypes to clarify ambiguous performance and architecture problems.
- Contribute visible technical work and, where appropriate, open-source contributions.
Requirements
- Deep understanding of modern serving frameworks and inference optimization techniques such as vLLM or TRT-LLM.
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding.
- Proficiency in C++, CUDA, Rust, or highly optimized Python, including profiling and optimizing NVIDIA GPU workloads.
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and handling thousands of concurrent connections.
- Public work such as substantial systems programming projects, major inference-engine open-source contributions, or deep-dive technical writing is useful.
- Ability to take a model from research through containerization, serving optimization, and reliable production deployment.
- PhD in computer science, physics, or mathematics, or equivalent practical experience building backend or machine learning systems.
- Ability to work independently in ambiguous environments and prioritize shipped, stable, high-impact engineering.
- Strong interest in understanding and improving core latency, throughput, architecture, and reliability challenges.
Benefits
- Relocation assistance is offered.
- The role is based in the Mountain View office with in-person collaboration.
- Equity and benefits are included.
Tech Stack
Categories
About Inworld
Inworld builds real-time voice AI and agent infrastructure for developers, including text-to-speech, speech-to-speech, speech recognition, and LLM routing delivered via APIs and SDKs. Customers use it to create interactive characters and agentic experiences in games, apps, and virtual worlds; revenue comes from usage-based and enterprise licensing. Founded in 2021 and headquartered in Mountain View, California, the company is privately held and works with game studios and large technology firms.
