
Member of Technical Staff - Training Platform
Prime Intellect2 months ago
Base Salary
$150k - $300k/yr
Responsibilities
- Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets.
- Build and maintain Helm charts and reproducible training stacks for trainers, inference servers, environment servers, and supporting services.
- Develop Python control-plane agents that monitor pods, report run state, and synchronize clusters with the platform.
- Implement scheduling and autoscaling for heterogeneous H100, H200, and B200 GPU hardware.
- Build node-local model caches, checkpoint pipelines, shared storage, and GitOps-based deployment workflows.
- Operate Prometheus, Grafana, Loki, and DCGM observability systems for GPU cluster debugging.
- Build hosted-training platform features including job submission, live run monitoring, logs, metrics, model and adapter management, and comparisons.
- Develop FastAPI backend services, REST APIs, real-time monitoring tools, streaming logs, step-level metrics, and failure-analysis capabilities.
- Ship product UI with Next.js, React, TypeScript, shadcn, Tailwind, tRPC, and TanStack Query.
- Interface with RL trainers, inference servers, and environment servers and productize new training capabilities.
Requirements
- Strong working knowledge of open model families, LoRA, QLoRA, full fine-tuning, RLHF, RLAIF, vLLM, SGLang, and TensorRT-LLM.
- Familiarity with H100, H200, and B200 GPU tradeoffs, NVLink, interconnects, and memory hierarchy.
- Understanding of distributed training concepts including data, tensor, pipeline, and expert parallelism, NCCL, and multi-node scheduling.
- Strong Kubernetes operations experience with Helm, CRDs, operators, KEDA, gang scheduling, and GPU Operator.
- Experience debugging production clusters using kubectl, pod lifecycle analysis, node troubleshooting, and networking.
- Cloud platform experience, preferably GCP, including GCS, GKE, Cloud Run, and Cloud Tasks.
- Experience with infrastructure automation using Helm, Terraform, and Ansible and with GitOps workflows.
- Experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM.
- Strong Python backend development experience with FastAPI, async programming, and SQLAlchemy.
- Ability to build Python control-plane agents that communicate with Kubernetes APIs.
- Comfort with TypeScript, React, Next.js, Tailwind, and shadcn for end-to-end product development.
- Experience designing REST and tRPC APIs and building developer tools, dashboards, and live-monitoring UIs.
- Experience in platform or infrastructure development, ideally both, and the ability to work across AI systems, infrastructure, and product development.
Benefits
- Cash compensation of $150K–$300K with significant equity.
- Flexible work arrangement with remote work or a San Francisco office option.
- Full visa sponsorship and relocation support.
- Professional development budget for courses and conferences.
- Regular team off-sites and conference attendance.
- Opportunity to contribute to open-source work and the broader AI community.
Tech Stack
AnsibleFastAPIGoogle Cloud PlatformGrafanaHelmKubernetesNext.jsPrometheusPythonReactTailwind CSSTerraformTypeScript