5 months ago
Remote, WorldwideMid Level / Senior / Staff+
H1B Sponsor
Responsibilities
- Build and maintain a Python fleet-tracking system covering server contracting, procurement, target use, pricing, availability, health, and RMAs.
- Develop server-management tooling for provisioning, health checks, GPU diagnostics, recovery, and alerting.
- Create and maintain metrics, dashboards, and alerts for GPU errors, disk failures, network issues, and thermal problems across the fleet.
- Use AI extensively to build tools and automate alerting and recovery.
- Implement OS security controls including hardening baselines, SELinux/AppArmor policies, SSH key management, vulnerability scanning, and compliance automation.
- Manage and optimize NVMe arrays, NFS, parallel file systems, and object storage for model weights, checkpoints, and ephemeral scratch data.
- Tune Linux systems for AI workloads, including kernel parameters, NUMA topology, CPU pinning, hugepages, I/O schedulers, and GPU driver stacks.
- Develop automated error-detection and recovery processes and work with partners to resolve issues.
Requirements
- At least 3 years of experience managing bare-metal and cloud-based server fleets at scale, including fleets of 100 or more nodes.
- Strong production software engineering skills in Python.
- Deep Linux systems knowledge covering boot processes, kernel tuning, networking, storage, systemd, cgroups, namespaces, and performance profiling.
- Strong experience with Ansible, Terraform, and cloud-init for configuration management and infrastructure as code.
- Solid understanding of LVM, RAID, NVMe, NFS, Lustre or GPFS, and Linux I/O stack tuning.
- Familiarity with hardware diagnostics and failure modes involving GPUs, NVMe, NICs, and memory.
- Experience building internal infrastructure-visibility tools or dashboards.
- Excellent communication skills and ability to drive technical decisions across teams.
- Nice-to-have experience with VLAN, VXLAN, ECMP, BGP, and tcpdump.
- Nice-to-have experience managing NVIDIA GPU infrastructure, including drivers, health monitoring, DCGM, NVLink/NVSwitch diagnostics, RDMA, and InfiniBand/RoCEv2.
- Nice-to-have experience with AMD GPUs, PXE/iPXE, Kickstart, libvirt, QEMU/KVM, SOC 2, or ISO 27001.
Benefits
- Work location is Turkey.
- Interesting and challenging work.
- Learning and growth opportunities.
- Regular team events and offsites.
