2 months ago
Remote, United StatesSenior / Staff+
H1B sponsor
Base Salary
$180k - $250k/yr
Responsibilities
- Design, automate, validate, and deliver the complete lifecycle of customer compute environments from provisioning through upgrades, recovery, and decommissioning.
- Use AI to automate and accelerate infrastructure delivery and operations.
- Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads.
- Build and maintain Linux images and automated operating-system provisioning workflows.
- Operate the NVIDIA GPU stack, including drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring.
- Design Kubernetes and data-center networking using Cilium, Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP.
- Configure distributed and shared storage for high-performance workloads.
- Build monitoring, alerting, diagnostics, and automated recovery for customer environments.
- Develop reusable tooling, standards, documentation, and runbooks.
- Collaborate with customers and internal teams to translate workload requirements into infrastructure designs.
Requirements
- At least 5 years of experience building and operating production Linux infrastructure.
- Strong production experience with Kubernetes on bare metal, including bootstrapping, upgrades, highly available control planes, etcd, containerd, CNI, CSI, ingress, load balancing, observability, security, and troubleshooting.
- Experience with Linux virtualization using KVM/QEMU, libvirt, and VFIO device passthrough.
- Experience operating NVIDIA GPUs on Linux and Kubernetes, including drivers, container runtimes, device plugins, GPU Operator, and GPU telemetry.
- Strong networking fundamentals covering TCP/IP, Layer 2/Layer 3, VLANs, routing, and packet-level troubleshooting with tcpdump and Wireshark.
- Practical scripting experience and experience with configuration-management tools such as Ansible.
- Ability to diagnose complex cross-layer infrastructure issues and drive technical decisions across teams.
- Production Slurm experience is preferred.
- Preferred experience includes NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX, hugepages, NUMA, CPU pinning, SR-IOV, and DPDK.
- Preferred distributed-storage experience includes Ceph, Lustre, or Weka.
- Preferred experience includes KubeVirt, OpenStack, IPsec, WireGuard, Tailscale, VXLAN, BGP, ECMP, BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init, NetBox, Nautobot, and Nornir.
- AI training, inference, or distributed GPU workload infrastructure experience is preferred.
- Python or Go proficiency is preferred.
Tech Stack
Categories
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
