
Member of Technical Staff (Software Engineer, GPU Cluster Infrastructure)
Perplexityabout 2 hours ago
Remote, United States +3 moreMid Level / Senior
H1B Sponsor
Base Salary
$250k - $485k/yr
Responsibilities
- Build a self-serve compute platform for inference engineers and researchers.
- Operate the GPU fleet, managing provisioning and lifecycle across providers.
- Develop scheduling and placement logic to optimize GPU resource usage.
- Support both long-running training jobs and production inference services.
- Manage Kubernetes orchestration for GPU workloads across multiple clusters.
- Implement fault tolerance, autoscaling, and observability for the GPU fleet.
- Set technical direction and collaborate with other engineering teams.
Requirements
- Deep experience with Kubernetes, including custom operators and multi-cluster federation.
- Experience managing GPU clusters at scale, including NVIDIA hardware and CUDA.
- Knowledge of orchestrating compute across multiple cloud providers.
- Strong fundamentals in distributed systems, including scheduling and fault tolerance.
- Proficiency in infrastructure and systems-level programming in Go, Rust, or C++.
- Experience supporting both training jobs and high-availability inference services.
- Ability to own problems end-to-end and navigate ambiguous situations.