GrepJob
Perplexity

Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)

Perplexity
Apply
about 2 hours ago
Remote, United States +3 moreMid Level / Senior
H1B Sponsor

Base Salary

$250k - $485k/yr

Responsibilities

  • Build a self-serve compute platform for inference engineers and researchers.
  • Operate the GPU fleet, managing provisioning and lifecycle across providers.
  • Develop scheduling and placement logic to optimize GPU resource usage.
  • Support both long-running training jobs and production inference services.
  • Manage Kubernetes orchestration for GPU workloads across multiple clusters.
  • Implement fault tolerance, autoscaling, and observability for the GPU fleet.
  • Set technical direction and collaborate with other engineering teams.

Requirements

  • Deep experience with Kubernetes, including custom operators and multi-cluster federation.
  • Experience managing GPU clusters at scale, including NVIDIA hardware and CUDA.
  • Knowledge of orchestrating compute across multiple cloud providers.
  • Strong fundamentals in distributed systems, including scheduling and fault tolerance.
  • Proficiency in infrastructure and systems-level programming in Go, Rust, or C++.
  • Experience supporting both training jobs and high-availability inference services.
  • Ability to own problems end-to-end and navigate ambiguous situations.

Categories

AI & MLData EngineeringDevOps