TensorWave

Staff Infrastructure Engineer – Kubernetes Platform

TensorWave
Apply
3 months ago
Las Vegas, NV, USAStaff+

Responsibilities

  • Design and evolve Kubernetes control plane architecture across regions
  • Define and implement multi-tenant cluster models, including shared control planes and virtual cluster approaches
  • Transition standalone clusters to regionally managed platform models
  • Define standards for isolation boundaries, resource segmentation, and policy enforcement
  • Own the reliability and production behavior of Kubernetes platforms
  • Participate in on-call rotation and lead incident response and root cause analysis
  • Diagnose and resolve control plane instability, API server saturation, scheduling issues, and resource contention
  • Ensure consistent cluster provisioning, upgrades, scaling, and lifecycle management
  • Design regional scaling and multi-data-center cluster deployment strategies
  • Define cluster topology and failure-domain strategies
  • Design cluster-level and regional ingress and egress architectures
  • Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, and CNI behavior
  • Improve observability across control plane components, cluster health, and performance
  • Define and implement resilience strategies aligned with platform goals
  • Collaborate with DevOps and infrastructure teams on automation, compute, storage, networking, and platform capabilities

Requirements

  • 7+ years of experience in infrastructure, platform engineering, or distributed systems
  • Deep experience operating Kubernetes at scale in production environments
  • Experience scaling Kubernetes across multiple clusters and multiple regions or data centers
  • Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd
  • Experience designing or evolving control plane architectures and multi-tenant cluster models
  • Strong Linux systems expertise and deep troubleshooting ability across Kubernetes, container runtimes, and networking stacks
  • Experience with CNI plugins, with Cilium preferred
  • Strong understanding of networking, traffic patterns, resource isolation, and scheduling
  • Experience in CSP, hyperscale, or equivalent large-scale environments is strongly preferred
  • Experience with virtual cluster technologies such as vcluster or Kamaji is preferred
  • Experience supporting GPU workloads in Kubernetes is preferred
  • Familiarity with NUMA-aware scheduling, topology-aware workloads, RDMA, high-throughput networking environments, and observability platforms such as Prometheus and Grafana is preferred

Benefits

  • Stock options
  • 100% paid medical, dental, and vision insurance for employees
  • Company Health Savings Account contributions
  • 100% paid short-term and long-term disability insurance for employees
  • Life and voluntary supplemental insurance options
  • Pet and legal insurance options
  • Supplementary health benefits, including discounted virtual healthcare appointments and serious illness support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid holidays
  • Parental leave
  • Other in-office perks

Tech Stack

GrafanaKubernetesLinuxPrometheus

Categories

DevOpsSite Reliability
TensorWave

About TensorWave

51-200 employees

TensorWave is an AMD-exclusive cloud platform designed specifically for AI workloads. Offering AMD Instinct MI325X and MI355X GPUs, TensorWave is a top-choice for training, fine-tuning, and inference. Visit tensorwave.com to learn more. Send us a message to try it for free.