fal

Senior/Staff Kubernetes Infrastructure Engineer

fal
Apply
2 hours ago
Remote, WorldwideSenior / Staff+
H1B Sponsor

Base Salary

$180k - $250k/yr

Responsibilities

  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning.
  • Use AI to automate and accelerate infrastructure delivery and operations.
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
  • Build and maintain Linux images and automated operating-system provisioning workflows.
  • Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
  • Design Kubernetes and data-center networking using Cilium, Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
  • Configure distributed and shared storage for high-performance workloads.
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
  • Develop reusable tooling, standards, documentation, and runbooks.
  • Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.

Requirements

  • At least 5 years of experience building and operating production Linux infrastructure.
  • Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
  • Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
  • Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
  • Strong networking fundamentals covering TCP/IP, Layer 2/Layer 3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
  • Practical scripting experience and experience with configuration-management tools such as Ansible.
  • Ability to diagnose complex cross-layer infrastructure issues and drive technical decisions across teams.
  • Production Slurm experience is preferred.
  • Preferred experience includes NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX, hugepages, NUMA, CPU pinning, SR-IOV, and DPDK.
  • Preferred distributed-storage experience includes Ceph, Lustre, or Weka.
  • Preferred experience includes KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, and Nornir.
  • AI training, inference, or distributed GPU workload infrastructure experience is preferred.
  • Python or Go proficiency is preferred.

Tech Stack

AnsibleGoKubernetesLinuxOpenStackPython

Categories

fal

About fal

51-200 employees

fal is a generative media platform that provides developers with access to the world's best generative image, video, and audio models through a unified API. Trusted by over 2.5 million developers and leading companies, fal offers the fastest inference engine for diffusion models, on-demand serverless GPUs, and dedicated compute clusters for frontier research. For press inquiries: press@fal.ai For customer support: support@fal.ai