12 months ago
Bengaluru, IndiaStaff+

Responsibilities

  • Design and evolve end-to-end infrastructure for ASR/TTS, LLM orchestration, Agentic RAG, and self-learning workflows.
  • Architect low-latency pipelines for real-time voice and chat AI with sub-second response targets.
  • Build elastic, multi-cloud distributed systems across AWS, GCP, and Azure.
  • Define and enforce latency, uptime, and throughput SLAs for AI services.
  • Drive observability, monitoring, resilience, and graceful failure strategies.
  • Optimize GPU and TPU utilization for cost-effective training and inference.
  • Embed security-by-design and implement controls for sensitive enterprise data and compliance requirements.
  • Translate AI research into production-grade platforms and mentor engineering teams on distributed systems and infrastructure design.
  • Evaluate inference optimization technologies including Triton, Riva, vLLM, and SGLang.

Requirements

  • 10–15 years of experience in large-scale systems architecture, including at least five years in principal architect-level roles.
  • Expertise in distributed systems, cloud-native architectures, and real-time pipelines.
  • Hands-on experience with containerization, Kubernetes orchestration, and microservices.
  • Strong background in scalable ML infrastructure, model serving, GPU/accelerator utilization, and ML CI/CD.
  • Experience architecting systems with low latency below 300ms, high throughput, and enterprise reliability.
  • Experience with conversational AI, speech systems, or real-time inference workloads.
  • Deep knowledge of MLOps platforms including Kubeflow, MLflow, Vertex AI, and SageMaker.
  • Familiarity with inference optimization frameworks including Triton, Nvidia Riva, vLLM, and SGLang.
  • Open-source contributions or patents in distributed systems, infrastructure, or ML tooling.
Nurix Therapeutics, Inc.

About Nurix Therapeutics, Inc.

201-500 employees
Contact me