
Member of Technical Staff (AI Infrastructure Engineer)
Perplexity4 months ago
London, United KingdomMid Level / Senior
H1B Sponsor
Responsibilities
- Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads.
- Manage and optimize Slurm-based HPC environments for distributed large-language-model training.
- Develop APIs and orchestration systems for training pipelines and inference services.
- Implement resource scheduling and job management across heterogeneous compute environments.
- Benchmark performance, diagnose bottlenecks, and improve training and inference infrastructure.
- Build ML-focused monitoring, alerting, and observability solutions for Kubernetes and Slurm workloads.
- Respond to system outages and collaborate across teams to maintain high uptime for training and inference services.
- Optimize cluster utilization and implement autoscaling for dynamic workload demands.
Requirements
- Strong expertise administering Kubernetes, including custom resource definitions, operators, and cluster management.
- Hands-on experience with Slurm workload management, job scheduling, resource allocation, and cluster optimization.
- Experience deploying and managing distributed training systems at scale, including GPU clusters and ML workloads.
- Strong understanding of container orchestration, distributed systems, networking, storage, and compute resource management.
- Proficiency in Python and C++ for systems and infrastructure automation.
- Hands-on experience with PyTorch in distributed training contexts.
- Experience developing APIs and managing distributed systems for batch and real-time workloads.
- Strong debugging, monitoring, and observability skills for containerized environments.
- Familiarity with LLM architecture and training processes, including Multi-Head Attention, Multi/Grouped-Query, and distributed training strategies.
- Preferred experience with Kubernetes operators, custom controllers, advanced Slurm federation and scheduling, GPU management, CUDA optimization, TensorFlow, distributed training libraries, HPC, parallel computing, high-performance networking, Terraform, Ansible, container registries, image optimization, and multi-stage builds.
- Demonstrated production experience managing large-scale Kubernetes deployments and Slurm clusters, with prior SRE, DevOps, or Platform Engineering experience focused on ML infrastructure.
- Experience supporting long-running training jobs and high-availability inference services; 3–5 years of relevant ML systems deployment experience is preferred.
Categories
About Perplexity
The most powerful answer engine. Powering curiosity with answers backed by up-to-date sources. This is where knowledge begins.