3 hours ago
Base Salary
$266k - $395k/yr
Responsibilities
- Design, build, and maintain scalable Kubernetes control-plane services, operators, and custom controllers.
- Develop Go and Python automation for end-to-end cluster provisioning, upgrades, patching, and deletion.
- Build GPU-aware orchestration for GPU scheduling and resource allocation.
- Partner with networking teams on CNI integration, high-performance fabrics, RDMA, and GPUDirect for AI workloads.
- Develop resilient distributed systems with timeouts, retries, backoff, and degraded-mode operation.
- Build inference platform services, including model-serving infrastructure, inference-load autoscaling, and multi-model deployment patterns.
- Create internal tools and CLIs that enable ML and AI teams to deploy and monitor inference services.
- Support and debug production systems through an on-call rotation.
Requirements
- 6+ years of software engineering experience with ownership of significant technical scope, such as driving projects from design through production or acting as a de facto technical lead.
- Deep understanding of Kubernetes internals, including controllers, schedulers, operators, CRDs, CSI, CNI, and Kubernetes extension patterns.
- Strong understanding of distributed systems fundamentals, fault tolerance, graceful degradation, and failure handling.
- Experience operating control planes and low-level components of large-scale Kubernetes clusters.
- Experience with observability at scale, including Prometheus, Grafana, distributed tracing, and actionable alerting.
- Strong programming skills in Go and Python and the ability to collaborate on shared codebases.
- Solid knowledge of Linux systems, networking, containers, and cloud infrastructure.
- Preferred: experience building or operating managed Kubernetes services such as GKE, EKS, or AKS, or working on Kubernetes control-plane components.
- Preferred: hands-on experience with NVIDIA GPU and networking technologies, including GPU Operator, device plugins, DCGM, MIG, Network Operator, and NCCL tuning.
- Preferred: familiarity with Slurm, KAI, Volcano, or Kueue and with GPU, InfiniBand, RDMA, high-performance computing, or AI/ML storage on Kubernetes.
- Contributions to CNCF projects or Kubernetes SIGs are a plus.
Benefits
- Hybrid work arrangement requiring presence in the San Francisco, San Jose, or Bellevue office 4 days per week, with Tuesday as the designated work-from-home day.
- Health, dental, and vision coverage for employees and dependents.
- Wellness and commuter stipends for select roles.
- 401(k) plan with a 2% company match for U.S. employees.
- Flexible paid time off plan.
- Cash and equity compensation are offered, with no specific compensation figures stated.
