Sandisk

AI SW Stack Deployment Architect (12+ years)

Sandisk
Apply
2 months ago
Bengaluru, IndiaStaff+

Responsibilities

  • Architect integration of vLLM, PyTorch, TensorFlow, and JAX/XLA into the accelerator stack.
  • Define framework, compiler, and runtime APIs and contracts.
  • Own LLM execution behavior, including batching, KV cache, and streaming inference.
  • Design and implement end-to-end deployment workflows for packaging, versioning, and reproducibility.
  • Drive performance optimization across models, frameworks, and runtimes.
  • Collaborate with compiler, runtime, and low-level software teams.
  • Support customer workloads, model onboarding, and debugging.

Requirements

  • 10+ years of experience in AI/ML systems or software architecture.
  • Strong experience with PyTorch, Transformers, and LLMs.
  • Hands-on experience with LLM deployment and scalable inference engine systems such as vLLM, Triton, or SGLang.
  • Experience building scalable AI platforms for cloud or edge environments.
  • Expertise in system design, APIs, and cross-layer integration.
  • Preferred: experience with vLLM or similar LLM serving systems, familiarity with XLA, MLIR, or compiler frameworks, exposure to GPU/NPU accelerators and runtime systems, and experience with distributed or multi-agent AI systems.

Benefits

  • Inclusive work environment with a stated commitment to diversity, belonging, accessibility, and accommodations.

Tech Stack

PyTorchTensorFlow

Categories

Sandisk

About Sandisk

5,001-10,000 employees
Contact me