7 months ago
Base Salary
$180k - $250k/yr
Responsibilities
- Build and maintain a Python fleet-tracking system covering server contracting, procurement, target use, pricing, availability, health, and RMAs.
- Develop tooling to automate server provisioning, health checks, GPU diagnostics, recovery, and alerting.
- Create and maintain metrics, dashboards, and alerts for GPU errors, disk failures, network issues, and thermal conditions.
- Use AI extensively to build tools and automate alerting and recovery.
- Implement OS security hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
- Manage and optimize NVMe arrays, NFS, parallel file systems, and object storage for model weights, checkpoints, and scratch data.
- Tune Linux systems for AI workloads, including kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver and container-runtime optimization.
- Develop automated error-detection and recovery processes.
- Work with partners to resolve technical issues.
Requirements
- At least three years of experience managing bare-metal and cloud-based server fleets at scale, including 100 or more nodes.
- Strong production software engineering skills in Python.
- Deep Linux systems knowledge, including boot processes, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
- Strong experience with Ansible, Terraform, and cloud-init for configuration management and infrastructure as code.
- Solid understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning.
- Familiarity with hardware diagnostics and failure modes involving GPUs, NVMe, NICs, and memory.
- Experience building internal infrastructure-visibility tools or dashboards.
- Excellent communication skills and ability to drive technical decisions across teams.
- Nice-to-have experience with VLAN, VXLAN, ECMP, BGP, and tcpdump.
- Nice-to-have experience managing NVIDIA GPU drivers, health monitoring, DCGM, NVLink/NVSwitch, RDMA, and InfiniBand/RoCEv2.
- Experience with AMD GPUs, PXE/iPXE, Kickstart, libvirt, QEMU/KVM, SOC 2, or ISO 27001 is a plus.
Benefits
- Relocation assistance to San Francisco.
- Health, dental, and vision insurance in the US.
- Regular team events and offsites.
- Position is located in San Francisco, California.
About fal
Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.
