Prime Intellect

Member of Technical Staff - Training Platform

Prime Intellect
Apply
2 months ago

Base Salary

$150k - $300k/yr

Responsibilities

  • Design and operate Kubernetes-based training and inference orchestration across multi-cluster, multi-cloud GPU fleets.
  • Build and maintain Helm charts and reproducible training stacks for trainers, inference servers, environment servers, and supporting services.
  • Develop Python control-plane agents that monitor pods, report run state, and synchronize clusters with the platform.
  • Implement scheduling and autoscaling for heterogeneous H100, H200, and B200 GPU hardware.
  • Build node-local model caches, checkpoint pipelines, shared storage, and GitOps-based deployment workflows.
  • Operate Prometheus, Grafana, Loki, and DCGM observability systems for GPU cluster debugging.
  • Build hosted-training platform features including job submission, live run monitoring, logs, metrics, model and adapter management, and comparisons.
  • Develop FastAPI backend services, REST APIs, real-time monitoring tools, streaming logs, step-level metrics, and failure-analysis capabilities.
  • Ship product UI with Next.js, React, TypeScript, shadcn, Tailwind, tRPC, and TanStack Query.
  • Interface with RL trainers, inference servers, and environment servers and productize new training capabilities.

Requirements

  • Strong working knowledge of open model families, LoRA, QLoRA, full fine-tuning, RLHF, RLAIF, vLLM, SGLang, and TensorRT-LLM.
  • Familiarity with H100, H200, and B200 GPU tradeoffs, NVLink, interconnects, and memory hierarchy.
  • Understanding of distributed training concepts including data, tensor, pipeline, and expert parallelism, NCCL, and multi-node scheduling.
  • Strong Kubernetes operations experience with Helm, CRDs, operators, KEDA, gang scheduling, and GPU Operator.
  • Experience debugging production clusters using kubectl, pod lifecycle analysis, node troubleshooting, and networking.
  • Cloud platform experience, preferably GCP, including GCS, GKE, Cloud Run, and Cloud Tasks.
  • Experience with infrastructure automation using Helm, Terraform, and Ansible and with GitOps workflows.
  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, and DCGM.
  • Strong Python backend development experience with FastAPI, async programming, and SQLAlchemy.
  • Ability to build Python control-plane agents that communicate with Kubernetes APIs.
  • Comfort with TypeScript, React, Next.js, Tailwind, and shadcn for end-to-end product development.
  • Experience designing REST and tRPC APIs and building developer tools, dashboards, and live-monitoring UIs.
  • Experience in platform or infrastructure development, ideally both, and the ability to work across AI systems, infrastructure, and product development.

Benefits

  • Cash compensation of $150K–$300K with significant equity.
  • Flexible work arrangement with remote work or a San Francisco office option.
  • Full visa sponsorship and relocation support.
  • Professional development budget for courses and conferences.
  • Regular team off-sites and conference attendance.
  • Opportunity to contribute to open-source work and the broader AI community.

Tech Stack

AnsibleFastAPIGoogle Cloud PlatformGrafanaHelmKubernetesNext.jsPrometheusPythonReactTailwind CSSTerraformTypeScript
Prime Intellect

About Prime Intellect

51-200 employees
Contact me