TensorWave

Staff Infrastructure Engineer – Kubernetes Platform

TensorWave
Apply
5 months ago
Las Vegas, NV, USAStaff+

Responsibilities

  • Design and evolve Kubernetes control plane architecture across regions
  • Define and implement multi-tenant cluster models, including shared control planes and virtual cluster approaches
  • Transition standalone clusters to regionally managed platform models
  • Define standards for isolation boundaries, resource segmentation, and policy enforcement
  • Own the reliability and production behavior of Kubernetes platforms
  • Participate in on-call rotation and lead incident response and root cause analysis
  • Diagnose and resolve control plane instability, API server saturation, scheduling issues, and resource contention
  • Ensure consistent cluster provisioning, upgrades, scaling, and lifecycle management
  • Design regional scaling and multi-data-center cluster deployment strategies
  • Define cluster topology and failure-domain strategies
  • Design cluster-level and regional ingress and egress architectures
  • Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, and CNI behavior
  • Improve observability across control plane components, cluster health, and performance
  • Define and implement resilience strategies aligned with platform goals
  • Collaborate with DevOps and infrastructure teams on automation, compute, storage, networking, and platform capabilities

Requirements

  • 7+ years of experience in infrastructure, platform engineering, or distributed systems
  • Deep experience operating Kubernetes at scale in production environments
  • Experience scaling Kubernetes across multiple clusters and multiple regions or data centers
  • Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd
  • Experience designing or evolving control plane architectures and multi-tenant cluster models
  • Strong Linux systems expertise and deep troubleshooting ability across Kubernetes, container runtimes, and networking stacks
  • Experience with CNI plugins, with Cilium preferred
  • Strong understanding of networking, traffic patterns, resource isolation, and scheduling
  • Experience in CSP, hyperscale, or equivalent large-scale environments is strongly preferred
  • Experience with virtual cluster technologies such as vcluster or Kamaji is preferred
  • Experience supporting GPU workloads in Kubernetes is preferred
  • Familiarity with NUMA-aware scheduling, topology-aware workloads, RDMA, high-throughput networking environments, and observability platforms such as Prometheus and Grafana is preferred

Benefits

  • Stock options
  • 100% paid medical, dental, and vision insurance for employees
  • Company Health Savings Account contributions
  • 100% paid short-term and long-term disability insurance for employees
  • Life and voluntary supplemental insurance options
  • Pet and legal insurance options
  • Supplementary health benefits, including discounted virtual healthcare appointments and serious illness support
  • Flexible Spending Account
  • 401(k)
  • Employee Assistance Program
  • Flexible PTO
  • Paid holidays
  • Parental leave
  • Other in-office perks

Tech Stack

GrafanaKubernetesLinuxPrometheus

Categories

DevOpsSite Reliability
TensorWave

About TensorWave

51-200 employees

TensorWave builds an AMD‑exclusive cloud platform for AI workloads, providing Instinct MI325X and MI355X GPU instances and tooling for training, fine‑tuning, and inference, including an inference engine. It sells infrastructure-as-a-service to startups and enterprises that need scalable AI compute. Founded in 2023 and headquartered in Las Vegas, Nevada, the company focuses on GPU‑accelerated cloud infrastructure for generative AI and machine learning teams.

Contact me