GrepJob
fal

Software Engineer, Infrastructure

fal
Apply
5 months ago
Remote, WorldwideMid Level / Senior / Staff+
H1B Sponsor

Responsibilities

  • Build and maintain a Python fleet-tracking system covering server contracting, procurement, target use, pricing, availability, health, and RMAs.
  • Develop server-management tooling for provisioning, health checks, GPU diagnostics, recovery, and alerting.
  • Create and maintain metrics, dashboards, and alerts for GPU errors, disk failures, network issues, and thermal problems across the fleet.
  • Use AI extensively to build tools and automate alerting and recovery.
  • Implement OS security controls including hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
  • Manage and optimize NVMe arrays, NFS, parallel file systems, and object storage for model weights, checkpoints, and ephemeral scratch data.
  • Tune Linux systems for AI workloads, including kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stacks.
  • Develop automated error-detection and recovery processes and work with partners to resolve issues.

Requirements

  • At least 3 years of experience managing bare-metal and cloud-based server fleets at scale, including fleets of 100 or more nodes.
  • Strong production software engineering skills in Python.
  • Deep Linux systems knowledge covering boot processes, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
  • Strong experience with Ansible, Terraform, and cloud-init for configuration management and infrastructure as code.
  • Solid understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning.
  • Familiarity with hardware diagnostics and failure modes involving GPUs, NVMe, NICs, and memory.
  • Experience building internal infrastructure-visibility tools or dashboards.
  • Excellent communication skills and ability to drive technical decisions across teams.
  • Nice-to-have experience with VLAN, VXLAN, ECMP, BGP, and tcpdump.
  • Nice-to-have experience managing NVIDIA GPU infrastructure, including drivers, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2.
  • Nice-to-have experience with AMD GPUs, PXE/iPXE, Kickstart, libvirt, QEMU/KVM, SOC 2, or ISO 27001.

Benefits

  • Work location is Turkey.
  • Interesting and challenging work.
  • Learning and growth opportunities.
  • Regular team events and offsites.

Tech Stack

Categories