
Member of Technical Staff - Inference
Prime Intellect2 months ago
Remote, United States or San Francisco, CA, USAMid Level
H1B sponsor
Base Salary
$150k - $300k/yr
Responsibilities
- Build a multi-tenant LLM serving platform across cloud GPU fleets.
- Design GPU-aware placement and scheduling algorithms for heterogeneous accelerators.
- Implement multi-region and multi-zone failover, traffic shifting, autoscaling, routing, and load balancing.
- Optimize model distribution and cold-start performance across clusters.
- Integrate and contribute to vLLM, SGLang, and TensorRT-LLM inference frameworks.
- Tune tensor, pipeline, and expert parallelism, prefix caching, memory management, quantization, and speculative decoding.
- Profile kernels, memory bandwidth, and transport, and develop reproducible latency and throughput performance suites.
- Embed and optimize distributed inference within the RL stack.
- Establish CI/CD with artifact promotion, performance gates, and reproducible builds.
- Build observability, incident-response, and SLO-management systems, and document architectures and playbooks.
Requirements
- At least 3 years of experience building and operating large-scale ML or LLM services with latency and availability SLOs.
- Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
- Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
- Deep understanding of prefill and decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism strategies.
- Ability to debug CUDA/NCCL, drivers and kernels, containers, service mesh and networking, and storage while owning incidents end to end.
- Experience with Python, PyTorch, AWS or GCP, Kubernetes, CUDA, NCCL, InfiniBand, GPU architecture, and GPU-aware scheduling.
- Preferred qualifications include CUDA or Triton kernel development, Nsight Systems or Nsight Compute profiling, Rust or C++, Kafka or PubSub, Redis, gRPC or Protobuf, Prometheus or Grafana, OpenTelemetry, Terraform, Ansible, and open-source infrastructure contributions.
Benefits
- Cash compensation is $150-300k with significant equity incentives.
- Flexible work arrangement with remote work or a San Francisco office option.
- Full visa sponsorship and relocation support.
- Professional development budget.
- Regular team off-sites and conference attendance.
- Opportunity to contribute to decentralized AI and RL through research and open-source work.
Tech Stack
AnsibleApache KafkaAWSC++Google Cloud PlatformGrafanagRPCKubernetesPrometheusPythonPyTorchRedisRustTerraform