2 days ago
Base Salary
$207k - $452k/yr
Responsibilities
- Design and build production-scale inference serving systems for transformer and multimodal models, including 100B+ and MoE architectures.
- Implement and tune speculative decoding, continuous batching, KV cache management, prefill/decode disaggregation, and INT4/INT8/FP8 quantization.
- Customize inference frameworks such as vLLM, TensorRT-LLM, and SGLang for production requirements.
- Write, optimize, and profile CUDA kernels and custom operations.
- Own end-to-end deployment, including model packaging, serving API design, latency SLO monitoring, and incident response.
- Partner with research teams to translate model architecture changes into inference-efficient implementations.
- Drive technical design and establish inference engineering practices across the team.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
- At least 5 years of software engineering experience, including significant production-depth experience with inference systems or ML infrastructure.
- Hands-on serving experience with vLLM, TensorRT-LLM, SGLang, or ONNX Runtime.
- GPU programming experience, including CUDA kernel development, memory optimization, and profiling with Nsight or equivalent tools.
- Production experience serving LLMs or large vision models, owning latency SLOs, debugging throughput regressions, and shipping performance optimizations.
- Depth in at least two of speculative decoding, continuous batching, KV cache design, quantization pipelines, and prefill/decode disaggregation.
- Strong systems programming ability in Python and C++, including the ability to read and modify framework internals.
- Preferred: Master’s or PhD in a relevant technical field.
- Preferred: experience with MoE models or 100B+ parameter deployments, disaggregated serving architectures, or multi-node inference.
- Preferred: compiler-level optimization experience with XLA, Triton, or similar technologies.
Benefits
- Structured hybrid work approach using office and remote work environments.
- Benefits and perks supporting physical, mental, emotional, and financial health, work-life balance, and community involvement.
- Applications are accepted through the anticipated position close date of October 16, 2026, with at least a five-day application window.
Categories
About Zoom
Zoom builds a cloud platform for video meetings, team chat, webinars, phone, rooms, and contact center used by businesses, schools, and public-sector organizations. It sells these collaboration and communications services as SaaS subscriptions with add-ons for events, telephony, and enterprise features, and offers developer APIs/SDKs. Founded in 2013 and headquartered in San Jose, California, Zoom Video Communications is a public company traded on Nasdaq.
