GrepJob
fal

Software Engineer, Infrastructure

fal
Apply
6 months ago
San Francisco, CA, USAMid Level
H1B Sponsor

Base Salary

$180k - $250k/yr

Responsibilities

  • Build and maintain a Python fleet-tracking system covering server contracting, procurement, target use, pricing, availability, health, RMAs, and lifecycle management.
  • Develop tooling to automate server provisioning, health checks, GPU diagnostics, recovery, and alerting.
  • Create and maintain metrics, dashboards, and alerts for GPU errors, disk failures, network issues, thermals, and other fleet hardware health conditions.
  • Use AI extensively to build tools and automate alerting and recovery.
  • Implement OS-level security through hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
  • Manage and optimize NVMe arrays, NFS, parallel file systems, and object storage for model weights, checkpoints, and ephemeral scratch data.
  • Tune Linux systems for AI workloads, including kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stacks.
  • Develop automated error-detection and recovery processes and work with partners to resolve issues that automation cannot fix.

Requirements

  • At least 3 years of experience managing bare-metal and cloud-based server fleets at scale, including fleets of 100 or more nodes.
  • Strong production software engineering experience in Python.
  • Deep Linux systems knowledge covering boot processes, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
  • Strong experience with configuration management and infrastructure-as-code using Ansible, Terraform, and cloud-init.
  • Solid understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning.
  • Familiarity with hardware diagnostics and failure modes involving GPUs, NVMe devices, NICs, and memory.
  • Experience building internal infrastructure visibility tools or dashboards.
  • Excellent communication skills and the ability to drive technical decisions across teams.
  • Nice-to-have experience with VLAN, VXLAN, ECMP, BGP, and tcpdump for network configuration and diagnostics.
  • Nice-to-have experience managing NVIDIA GPU infrastructure, including drivers, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2.
  • Experience with AMD GPUs, PXE/iPXE, Kickstart, libvirt, QEMU/KVM, and bare-metal or VM provisioning is also valued.
  • Familiarity with SOC 2 and ISO 27001 compliance frameworks is preferred.

Benefits

  • Equity and benefits are offered in addition to base compensation.
  • Health, dental, and vision insurance is available in the US.
  • Relocation assistance to San Francisco is offered.
  • Regular team events and offsites are provided.
  • The role is located in San Francisco, California.

Tech Stack

Categories