Nutanix

Senior Software Engineer - LLM Inference

Nutanix
Apply
2 hours ago
Vancouver, CanadaSenior

Responsibilities

  • Architect, design, and develop horizontally scalable, containerized, fault-tolerant services on Kubernetes for enterprise AI and LLM workloads.
  • Build and operate high-performance inference and platform services delivering low-latency, high-throughput experiences for Generative AI and Agentic AI applications.
  • Design and optimize distributed systems, storage, networking, and low-level infrastructure components.
  • Develop multi-tenant platform services for on-premises, hybrid, and cloud-based AI deployments.
  • Implement observability architectures using Prometheus, Grafana, Datadog, OpenTelemetry, and related cloud-native monitoring tools.
  • Debug production issues, perform root-cause analysis, and improve reliability, resiliency, and operational efficiency.
  • Build and maintain CI/CD pipelines and deployment automation.
  • Implement LLM serving capabilities including request routing, rate limiting, token streaming, load balancing, quota management, and usage budgeting.
  • Contribute across architecture, design, development, testing, experimentation, performance analysis, deployment, and operations.
  • Collaborate with product management, AI, and software engineering teams; review code and design documents; and help shape the technical direction of the Enterprise AI Platform.

Requirements

  • 8+ years of experience developing maintainable, modular, resilient, fail-safe, and long-lived software products.
  • Strong computer science fundamentals in data structures, algorithms, operating systems, networking, and distributed systems.
  • Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
  • Production backend development experience using Go, Python, C++, or Rust.
  • Experience owning CI/CD pipelines and release automation end-to-end.
  • Strong understanding of datacenter architecture, including compute, storage, networking, and virtualization.
  • Experience deploying software across on-premises, cloud, and hybrid environments.
  • Experience designing and tuning high-performance, performance-sensitive system software.
  • Understanding of distributed computing, distributed data stores, and large-scale service architectures.
  • Experience diagnosing production performance issues with observability and monitoring platforms.
  • Familiarity with LLM serving concepts such as rate limiting, token streaming, request scheduling, load balancing, quota management, and usage budgeting.
  • Familiarity with LLM concepts including reasoning workflows, tool calling, prompt templates, and agents.
  • Experience building multi-tenant services on virtualized or containerized infrastructure.
  • Master's degree in Computer Science or equivalent practical experience.
  • Preferred experience with PyTorch, TensorFlow, GPU-based systems, vLLM, DeepSpeed, Hugging Face TGI, Triton, RAG, vector databases, AI orchestration frameworks, open-source contributions, production AI platforms, LLM APIs, agentic systems, inference infrastructure, SSE/WebSockets, and prompt guardrails.

Benefits

  • Hybrid work arrangement with a minimum of 3 onsite days per week in applicable locations, including Vancouver.
  • RRSP with dollar-for-dollar matching up to 7% of base salary.
  • Mental health coverage and paramedical benefits.
  • Fully paid maternity and parental leave and generous bereavement leave, including pet loss.
  • RSUs and an Employee Stock Purchase Plan with a 15% discount.
  • The posting states that the application deadline is 40 days from the posting date and may be removed earlier if the role is filled.

Tech Stack

C++DatadogDockerGoGrafanaKubernetesPrometheusPythonPyTorchRustTensorFlow
Nutanix

About Nutanix

5,001-10,000 employees

Nutanix is a global leader in cloud software, offering organizations a single platform for running apps and data across clouds. With Nutanix, companies can reduce complexity and simplify operations, freeing them to focus on their business outcomes. Building on its legacy as the pioneer of hyperconverged infrastructure, Nutanix is trusted by companies worldwide to power hybrid multicloud environments consistently, simply, and cost-effectively. Learn more at www.nutanix.com or follow us on social media @nutanix.

Contact me