DRW

AI Inference Platform Engineer

DRW
Apply
5 hours ago

Base Salary

$200k - $250k/yr

Responsibilities

  • Optimize LLM inference performance across NVIDIA GPU architectures and inference runtimes.
  • Build performance profiling and observability from individual GPU kernels through multi-node inference systems.
  • Design and optimize KV cache, distributed inference, caching, routing, memory tiering, and prefill/decode architectures.
  • Onboard new models by selecting runtime, precision, sharding, memory, batching, cache, and serving configurations.
  • Maintain validated performance profiles and regression testing for model and hardware combinations.
  • Measure quality equivalence across serving configurations, including KV cache quantization, speculative decoding, precision choices, and model routing.
  • Manage model-serving versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
  • Partner with SRE and platform teams on deployment automation, model distribution, observability, production readiness, and reliable operations.
  • Optimize model placement, scaling, resource allocation, utilization, and cost across the inference fleet.
  • Design and operate multi-tenant GPU scheduling and workload isolation while balancing latency SLOs, throughput, and workload priority.

Requirements

  • Hands-on experience serving LLMs on NVIDIA GPUs and familiarity with Hopper, Blackwell, HBM, Tensor Cores, NVLink, and NVSwitch.
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.
  • Practical knowledge of continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching.
  • Experience evaluating model quality equivalence with evaluation harnesses, task-specific benchmarks, and regression detection.
  • Experience designing and tuning distributed inference systems with tensor parallelism, multi-node deployments, and disaggregated prefill and decode.
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS.
  • Proficiency with GPU performance and observability tools including Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux and systems-performance fundamentals across hardware, drivers, runtimes, networking, and application layers.
  • Production experience with model-serving infrastructure, CI/CD, automated testing, observability, and production readiness.
  • Measurement-driven performance optimization, ownership across hardware/runtime/model/infrastructure boundaries, clear communication, and the ability to evaluate unfamiliar models, runtimes, and hardware.

Benefits

  • Annual discretionary bonus eligibility.
  • Group medical, pharmacy, dental, and vision insurance.
  • 401k with discretionary employer match.
  • Short- and long-term disability insurance, life and AD&D insurance, health savings accounts, and flexible spending accounts.

Tech Stack

GrafanaLinuxPrometheus
DRW

About DRW

1,001-5,000 employees

DRW is a privately held proprietary trading firm that uses its own capital to trade and invest across global markets and asset classes, powered by technology and research. Founded in 1992 and headquartered in Chicago, it develops trading systems for strategies ranging from seconds to years and manages its own risk rather than external client money. The firm operates internationally from multiple offices.

Contact me