fal

Software Engineer, Infrastructure

fal
Apply
6 months ago
Remote, TurkeyMid Level / Senior / Staff+

Responsibilities

  • Build and maintain a Python fleet-tracking system covering server contracting, procurement, intended use, pricing, availability, health, and RMAs.
  • Develop tooling to automate server provisioning, health checks, GPU diagnostics, recovery, and alerting.
  • Create and maintain metrics, dashboards, and alerts for GPU errors, disk failures, network issues, and thermal problems.
  • Implement OS-level security through hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
  • Manage and optimize NVMe arrays, NFS, parallel file systems, and object storage for model weights, checkpoints, and ephemeral scratch data.
  • Tune Linux systems for AI workloads, including kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stacks.
  • Develop automated error detection and recovery processes.
  • Work with partners to resolve technical issues and drive recovery when automation cannot resolve failures.
  • Use AI extensively to build tools and automate alerting and recovery.

Requirements

  • At least 3 years of experience managing bare-metal and cloud-based server fleets at scale, including 100 or more nodes.
  • Strong production software engineering skills in Python.
  • Deep Linux systems knowledge, including boot processes, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
  • Strong experience with Ansible, Terraform, and cloud-init for configuration management and infrastructure as code.
  • Understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning.
  • Familiarity with hardware diagnostics and failure modes involving GPUs, NVMe, NICs, and memory.
  • Experience building internal infrastructure visibility tools or dashboards.
  • Excellent communication skills and ability to drive technical decisions across teams.
  • Nice-to-have experience with VLAN, VXLAN, ECMP, BGP, and tcpdump.
  • Nice-to-have experience with NVIDIA GPU infrastructure, including driver management, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2.
  • Nice-to-have experience with AMD GPUs, PXE/iPXE, Kickstart, libvirt, QEMU/KVM, SOC 2, or ISO 27001.

Benefits

  • Work location is Turkey.
  • Learning and growth opportunities.
  • Regular team events and offsites.

Tech Stack

Categories

fal

About fal

51-200 employees

Fal builds a generative media platform that gives developers a unified API to run state-of-the-art image, video, and audio models. It provides serverless GPUs, high-performance inference, and dedicated compute clusters so teams can customize, deploy, and scale models in production. The company is privately held and headquartered in San Francisco, serving both startups and enterprises through a commercial API and managed infrastructure.

Contact me