6 months ago
Zürich, SwitzerlandStaff+
Responsibilities
- Take models from the research team through containerization, serving optimization, and reliable production operation.
- Optimize inference latency and throughput using model acceleration and serving techniques.
- Build and operate high-performance systems using GPU and distributed inference infrastructure.
- Design and improve scaling, load balancing, and multi-GPU or multi-node serving for thousands of concurrent connections.
- Own performance, reliability, and stability as first-class production requirements.
- Investigate ambiguous systems problems through benchmarks and prototypes and drive solutions through deployment.
Requirements
- Deep understanding of modern serving frameworks and inference optimization techniques such as vLLM or TRT-LLM.
- Hands-on experience with quantization, distillation, caching strategies, continuous batching, paged attention, and speculative decoding.
- Proficiency in C++, CUDA, Rust, or highly optimized Python, including profiling code and optimizing performance on NVIDIA GPUs.
- Experience with Kubernetes, Ray, custom load balancing, multi-GPU or multi-node inference, and handling thousands of concurrent connections reliably.
- Public evidence of work such as substantial systems-programming projects, open-source contributions to major inference engines, or detailed technical write-ups.
- Ability to take a model from research through containerization, serving optimization, and reliable production deployment.
- PhD in computer science, physics, or mathematics, or equivalent practical experience building backend or machine-learning systems.
- Professional fluency in written and spoken English.
- Candidates must already have the legal right to work in Switzerland; visa sponsorship is not available.
Benefits
- Remote within Switzerland
- Full-time, permanent employment
- Employment via Employer of Record (EOR)
- Potential future U.S. visa and relocation support for relocation to the San Francisco Bay Area, subject to business needs and work authorization requirements
Tech Stack
Categories
About Inworld
Inworld builds real-time voice AI and agent infrastructure for developers, including text-to-speech, speech-to-speech, speech recognition, and LLM routing delivered via APIs and SDKs. Customers use it to create interactive characters and agentic experiences in games, apps, and virtual worlds; revenue comes from usage-based and enterprise licensing. Founded in 2021 and headquartered in Mountain View, California, the company is privately held and works with game studios and large technology firms.
