Prime Intellect

Member of Technical Staff - Inference

Prime Intellect
Apply
2 months ago
Remote, United States or San Francisco, CA, USAMid Level
H1B sponsor

Base Salary

$150k - $300k/yr

Responsibilities

  • Build a multi-tenant LLM serving platform across cloud GPU fleets.
  • Design GPU-aware placement and scheduling algorithms for heterogeneous accelerators.
  • Implement multi-region and multi-zone failover, traffic shifting, autoscaling, routing, and load balancing.
  • Optimize model distribution and cold-start performance across clusters.
  • Integrate and contribute to vLLM, SGLang, and TensorRT-LLM inference frameworks.
  • Tune tensor, pipeline, and expert parallelism, prefix caching, memory management, quantization, and speculative decoding.
  • Profile kernels, memory bandwidth, and transport, and develop reproducible latency and throughput performance suites.
  • Embed and optimize distributed inference within the RL stack.
  • Establish CI/CD with artifact promotion, performance gates, and reproducible builds.
  • Build observability, incident-response, and SLO-management systems, and document architectures and playbooks.

Requirements

  • At least 3 years of experience building and operating large-scale ML or LLM services with latency and availability SLOs.
  • Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
  • Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
  • Deep understanding of prefill and decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism strategies.
  • Ability to debug CUDA/NCCL, drivers and kernels, containers, service mesh and networking, and storage while owning incidents end to end.
  • Experience with Python, PyTorch, AWS or GCP, Kubernetes, CUDA, NCCL, InfiniBand, GPU architecture, and GPU-aware scheduling.
  • Preferred qualifications include CUDA or Triton kernel development, Nsight Systems or Nsight Compute profiling, Rust or C++, Kafka or PubSub, Redis, gRPC or Protobuf, Prometheus or Grafana, OpenTelemetry, Terraform, Ansible, and open-source infrastructure contributions.

Benefits

  • Cash compensation is $150-300k with significant equity incentives.
  • Flexible work arrangement with remote work or a San Francisco office option.
  • Full visa sponsorship and relocation support.
  • Professional development budget.
  • Regular team off-sites and conference attendance.
  • Opportunity to contribute to decentralized AI and RL through research and open-source work.
Prime Intellect

About Prime Intellect

51-200 employees
Contact me