2 hours ago
Remote, WorldwideSenior / Staff+
H1B Sponsor
Base Salary
$180k - $250k/yr
Responsibilities
- Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning.
- Use AI to automate and accelerate infrastructure delivery and operations.
- Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
- Build and maintain Linux images and automated operating-system provisioning workflows.
- Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
- Design Kubernetes and data-center networking using Cilium, Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
- Configure distributed and shared storage for high-performance workloads.
- Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
- Develop reusable tooling, standards, documentation, and runbooks.
- Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.
Requirements
- At least 5 years of experience building and operating production Linux infrastructure.
- Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
- Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
- Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
- Strong networking fundamentals covering TCP/IP, Layer 2/Layer 3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
- Practical scripting experience and experience with configuration-management tools such as Ansible.
- Ability to diagnose complex cross-layer infrastructure issues and drive technical decisions across teams.
- Production Slurm experience is preferred.
- Preferred experience includes NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX, hugepages, NUMA, CPU pinning, SR-IOV, and DPDK.
- Preferred distributed-storage experience includes Ceph, Lustre, or Weka.
- Preferred experience includes KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, and Nornir.
- AI training, inference, or distributed GPU workload infrastructure experience is preferred.
- Python or Go proficiency is preferred.
Tech Stack
Categories
About fal
fal is a generative media platform that provides developers with access to the world's best generative image, video, and audio models through a unified API. Trusted by over 2.5 million developers and leading companies, fal offers the fastest inference engine for diffusion models, on-demand serverless GPUs, and dedicated compute clusters for frontier research. For press inquiries: press@fal.ai For customer support: support@fal.ai
