2 months ago
Bengaluru, IndiaStaff+
Responsibilities
- Architect integration of vLLM, PyTorch, TensorFlow, and JAX/XLA into the accelerator stack.
- Define framework, compiler, and runtime APIs and contracts.
- Own LLM execution behavior, including batching, KV cache, and streaming inference.
- Design and implement end-to-end deployment workflows for packaging, versioning, and reproducibility.
- Drive performance optimization across models, frameworks, and runtimes.
- Collaborate with compiler, runtime, and low-level software teams.
- Support customer workloads, model onboarding, and debugging.
Requirements
- 10+ years of experience in AI/ML systems or software architecture.
- Strong experience with PyTorch, Transformers, and LLMs.
- Hands-on experience with LLM deployment and scalable inference engine systems such as vLLM, Triton, or SGLang.
- Experience building scalable AI platforms for cloud or edge environments.
- Expertise in system design, APIs, and cross-layer integration.
- Preferred: experience with vLLM or similar LLM serving systems, familiarity with XLA, MLIR, or compiler frameworks, exposure to GPU/NPU accelerators and runtime systems, and experience with distributed or multi-agent AI systems.
Benefits
- Inclusive work environment with a stated commitment to diversity, belonging, accessibility, and accommodations.
Tech Stack
PyTorchTensorFlow