Perplexity

Member of Technical Staff (AI Infrastructure Engineer)

Perplexity
Apply
4 months ago
London, United KingdomMid Level / Senior
H1B Sponsor

Responsibilities

  • Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads.
  • Manage and optimize Slurm-based HPC environments for distributed large-language-model training.
  • Develop APIs and orchestration systems for training pipelines and inference services.
  • Implement resource scheduling and job management across heterogeneous compute environments.
  • Benchmark performance, diagnose bottlenecks, and improve training and inference infrastructure.
  • Build ML-focused monitoring, alerting, and observability solutions for Kubernetes and Slurm workloads.
  • Respond to system outages and collaborate across teams to maintain high uptime for training and inference services.
  • Optimize cluster utilization and implement autoscaling for dynamic workload demands.

Requirements

  • Strong expertise administering Kubernetes, including custom resource definitions, operators, and cluster management.
  • Hands-on experience with Slurm workload management, job scheduling, resource allocation, and cluster optimization.
  • Experience deploying and managing distributed training systems at scale, including GPU clusters and ML workloads.
  • Strong understanding of container orchestration, distributed systems, networking, storage, and compute resource management.
  • Proficiency in Python and C++ for systems and infrastructure automation.
  • Hands-on experience with PyTorch in distributed training contexts.
  • Experience developing APIs and managing distributed systems for batch and real-time workloads.
  • Strong debugging, monitoring, and observability skills for containerized environments.
  • Familiarity with LLM architecture and training processes, including Multi-Head Attention, Multi/Grouped-Query, and distributed training strategies.
  • Preferred experience with Kubernetes operators, custom controllers, advanced Slurm federation and scheduling, GPU management, CUDA optimization, TensorFlow, distributed training libraries, HPC, parallel computing, high-performance networking, Terraform, Ansible, container registries, image optimization, and multi-stage builds.
  • Demonstrated production experience managing large-scale Kubernetes deployments and Slurm clusters, with prior SRE, DevOps, or Platform Engineering experience focused on ML infrastructure.
  • Experience supporting long-running training jobs and high-availability inference services; 3–5 years of relevant ML systems deployment experience is preferred.
Perplexity

About Perplexity

201-500 employees

The most powerful answer engine. Powering curiosity with answers backed by up-to-date sources. This is where knowledge begins.