
Staff Infrastructure Engineer – Kubernetes Platform
TensorWave3 months ago
Las Vegas, NV, USAStaff+
Responsibilities
- Design and evolve Kubernetes control plane architecture across regions
- Define and implement multi-tenant cluster models, including shared control planes and virtual cluster approaches
- Transition standalone clusters to regionally managed platform models
- Define standards for isolation boundaries, resource segmentation, and policy enforcement
- Own the reliability and production behavior of Kubernetes platforms
- Participate in on-call rotation and lead incident response and root cause analysis
- Diagnose and resolve control plane instability, API server saturation, scheduling issues, and resource contention
- Ensure consistent cluster provisioning, upgrades, scaling, and lifecycle management
- Design regional scaling and multi-data-center cluster deployment strategies
- Define cluster topology and failure-domain strategies
- Design cluster-level and regional ingress and egress architectures
- Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, and CNI behavior
- Improve observability across control plane components, cluster health, and performance
- Define and implement resilience strategies aligned with platform goals
- Collaborate with DevOps and infrastructure teams on automation, compute, storage, networking, and platform capabilities
Requirements
- 7+ years of experience in infrastructure, platform engineering, or distributed systems
- Deep experience operating Kubernetes at scale in production environments
- Experience scaling Kubernetes across multiple clusters and multiple regions or data centers
- Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd
- Experience designing or evolving control plane architectures and multi-tenant cluster models
- Strong Linux systems expertise and deep troubleshooting ability across Kubernetes, container runtimes, and networking stacks
- Experience with CNI plugins, with Cilium preferred
- Strong understanding of networking, traffic patterns, resource isolation, and scheduling
- Experience in CSP, hyperscale, or equivalent large-scale environments is strongly preferred
- Experience with virtual cluster technologies such as vcluster or Kamaji is preferred
- Experience supporting GPU workloads in Kubernetes is preferred
- Familiarity with NUMA-aware scheduling, topology-aware workloads, RDMA, high-throughput networking environments, and observability platforms such as Prometheus and Grafana is preferred
Benefits
- Stock options
- 100% paid medical, dental, and vision insurance for employees
- Company Health Savings Account contributions
- 100% paid short-term and long-term disability insurance for employees
- Life and voluntary supplemental insurance options
- Pet and legal insurance options
- Supplementary health benefits, including discounted virtual healthcare appointments and serious illness support
- Flexible Spending Account
- 401(k)
- Employee Assistance Program
- Flexible PTO
- Paid holidays
- Parental leave
- Other in-office perks
Tech Stack
Categories
DevOpsSite Reliability
About TensorWave
TensorWave is an AMD-exclusive cloud platform designed specifically for AI workloads. Offering AMD Instinct MI325X and MI355X GPUs, TensorWave is a top-choice for training, fine-tuning, and inference. Visit tensorwave.com to learn more. Send us a message to try it for free.