Nutanix

Staff Engineer - AI/Inference /Gateway

Nutanix
Apply
3 hours ago
Vancouver, CanadaStaff+

Responsibilities

  • Architect, design, and develop scalable, containerized, fault-tolerant services on Kubernetes for enterprise AI and LLM workloads.
  • Build and operate high-performance inference and platform services for Generative AI and agentic AI applications.
  • Design distributed systems, storage, networking, low-level infrastructure, and multi-tenant services for on-premises, hybrid, and cloud deployments.
  • Implement LLM serving capabilities including request routing, rate limiting, token streaming, load balancing, quota management, and usage budgeting.
  • Design observability architectures and diagnose production issues using monitoring and observability platforms.
  • Build and maintain CI/CD pipelines, release automation, and deployment automation.
  • Improve platform reliability, resiliency, performance, and operational efficiency through root-cause analysis and performance tuning.
  • Collaborate with product management, AI, and software engineering teams across the product lifecycle.
  • Review code and design documents, provide product feedback, contribute to open-source projects, and help shape technical direction.

Requirements

  • 8+ years of experience developing maintainable, modular, resilient, fail-safe, and long-lived software products.
  • Strong fundamentals in data structures, algorithms, operating systems, networking, and distributed systems.
  • Hands-on experience with Docker, Kubernetes, and cloud-native architectures.
  • Production backend development experience with Go, Python, C++, or Rust.
  • Experience owning CI/CD pipelines and release automation end-to-end.
  • Strong understanding of datacenter architecture, including compute, storage, networking, and virtualization.
  • Experience deploying software across on-premises, cloud, and hybrid environments.
  • Experience designing and tuning high-performance system software and diagnosing production performance issues.
  • Understanding of distributed computing, distributed data stores, and large-scale service architectures.
  • Familiarity with LLM serving concepts, reasoning workflows, tool calling, prompt templates, and agents.
  • Experience building multi-tenant services on virtualized or containerized infrastructure.
  • Master’s degree in Computer Science or equivalent practical experience.
  • Preferred experience includes PyTorch, TensorFlow, GPU acceleration, vLLM, DeepSpeed, Hugging Face TGI, Triton, RAG, vector databases, AI orchestration frameworks, LLM APIs, agentic systems, inference infrastructure, SSE, WebSockets, prompt guardrails, and open-source contributions.

Benefits

  • Hybrid work arrangement with employees in applicable locations expected to work onsite at least 3 days per week.
  • RRSP with dollar-for-dollar matching up to 7% of base salary.
  • Mental health coverage and paramedical benefits.
  • Fully paid maternity and parental leave and generous bereavement leave.
  • RSUs and an Employee Stock Purchase Plan with a 15% discount.
  • Role location is Vancouver, Canada, with an annual base pay range of CAD $171,000 to CAD $257,000.
  • Application deadline is 40 days from the posting date, though the posting may be removed earlier if filled.

Tech Stack

C++DatadogDockerGoGrafanaKubernetesPrometheusPythonPyTorchRustTensorFlow
Nutanix

About Nutanix

5,001-10,000 employees

Nutanix is a global leader in cloud software, offering organizations a single platform for running apps and data across clouds. With Nutanix, companies can reduce complexity and simplify operations, freeing them to focus on their business outcomes. Building on its legacy as the pioneer of hyperconverged infrastructure, Nutanix is trusted by companies worldwide to power hybrid multicloud environments consistently, simply, and cost-effectively. Learn more at www.nutanix.com or follow us on social media @nutanix.

Contact me