6 months ago
London, United KingdomStaff+
Responsibilities
- Take models from the research team through containerization, serving optimization, and reliable production deployment
- Optimize inference latency and throughput using modern serving frameworks and model acceleration techniques
- Profile and improve high-performance C++, CUDA, Rust, or Python systems running on NVIDIA GPUs
- Design and operate distributed, multi-GPU and multi-node inference systems capable of handling thousands of concurrent connections
- Evaluate solutions through benchmarks and prototypes and make architecture decisions based on performance and reliability
- Own production stability and treat performance, latency, and reliability as first-class product features
- Share impactful technical work and open-source contributions where appropriate
Requirements
- Deep understanding of modern serving frameworks and inference optimization techniques such as vLLM or TRT-LLM
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding
- Proficiency in C++, CUDA, Rust, or highly optimized Python, including profiling and optimizing NVIDIA GPU workloads
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU and multi-node inference, and handling thousands of concurrent connections
- Non-trivial systems programming projects, major open-source inference engine contributions, or deep technical write-ups are useful evidence of experience
- Ability to take a model from research through containerization, serving optimization, and reliable production operation
- PhD in computer science, physics, mathematics, or equivalent practical experience building backend or machine learning systems
- Strong ownership, learning agility, and ability to work effectively with ambiguous problems
Benefits
- Equity and benefits in addition to base pay
- Full-time position in the United Kingdom
- Candidates must already have the legal right to work in the UK; visa sponsorship is unavailable
- Potential future U.S. visa and relocation support for relocation to the San Francisco Bay Area, subject to business needs and work authorization requirements
- Flat structure, fast iterations, minimal process theater, and support for impactful open-source contributions
Tech Stack
Categories
About Inworld
Inworld builds real-time voice AI and agent infrastructure for developers, including text-to-speech, speech-to-speech, speech recognition, and LLM routing delivered via APIs and SDKs. Customers use it to create interactive characters and agentic experiences in games, apps, and virtual worlds; revenue comes from usage-based and enterprise licensing. Founded in 2021 and headquartered in Mountain View, California, the company is privately held and works with game studios and large technology firms.
