Biohub

Staff AI Infrastructure Engineer

Biohub
Apply
4 months ago
Foster City, CA, USAStaff+
H1B Sponsor

Base Salary

$241k - $331k/yr

Responsibilities

  • Own reliability, observability, and incident response for multi-site GPU clusters running Slurm on Kubernetes.
  • Debug and resolve infrastructure failures across storage, networking, scheduling, and GPU compute layers.
  • Design and execute GPU cluster scaling plans while validating storage, networking, interconnect, and scheduler behavior.
  • Build automation and tooling for capacity planning, GPU utilization monitoring, workload manager policy management, and pod lifecycle automation.
  • Drive configuration-as-code practices through reproducible, auditable, version-controlled pipelines.
  • Collaborate with AI researchers and hero-run leads to design infrastructure for frontier-scale training workloads.
  • Own technical vendor relationships, including SEV1 escalation, partner coordination, root-cause analysis, and SLA accountability.
  • Project GPU demand, manage expansion across GPU generations, and coordinate multi-cluster strategy.
  • Improve operational resilience by reducing detection and resolution times, automating toil, and developing scalable runbooks.

Requirements

  • 8+ years of AI/ML infrastructure engineering experience with deep expertise in HPC/Slurm operations, Kubernetes at scale, distributed systems debugging, or GPU compute infrastructure.
  • Strong Linux systems fundamentals, including TCP/IP, InfiniBand, RDMA, MTU/MSS/PMTUD, NFS, VAST, WEKA, POSIX semantics, cgroups, namespaces, eBPF, and sysctls.
  • Hands-on Kubernetes and cloud-native infrastructure experience, including pod lifecycle management, CNI plugins, StatefulSets, Helm, ArgoCD, or equivalent GitOps tooling; Cilium is preferred.
  • Experience with HPC workload managers, with Slurm strongly preferred; QoS, partitions, preemption, accounting, and Sunk/CoreWeave patterns are a plus.
  • Proficiency in Python and Bash; Go, Rust, or C/C++ experience is a plus.
  • Experience with observability stacks such as Prometheus/VictoriaMetrics, Grafana, DCGM metrics, and distributed tracing.
  • Ability to debug complex multi-system failures, form hypotheses quickly, design controlled experiments, and identify root causes under pressure.
  • Strong written and verbal communication skills for incident summaries, vendor escalations, and system design documentation.
  • Experience with distributed AI training infrastructure, including NCCL, PyTorch DDP, multi-node job debugging, checkpoint/restart patterns, or large-scale training container environments, is a bonus.

Benefits

  • Hybrid role requiring onsite work at least 60% of the working month, approximately 3 days per week, with schedule set by the manager.
  • Generous employer match on 401(k) contributions.
  • Paid time off for volunteering.
  • Funding for select family-forming benefits.
  • Relocation support.

Tech Stack

Categories

DevOpsSite Reliability
Biohub

About Biohub

201-500 employees

Our mission is to help scientists cure or prevent all disease. At Biohub, we build the technology to help scientists around the world use AI-powered biology to study how cells operate, organize, and work as part of systems to understand why disease happens and how to correct it. With unprecedented scale of compute, AI research and engineering, and state-of-the-art technology for measuring, imaging, and programming biology, Biohub is leading the first large-scale scientific initiative to push the frontier of artificial intelligence for biology.