11 months ago
Responsibilities
- Optimize CUDA and Triton GPU kernels, FlashAttention, paged attention, and CUDA Graphs.
- Improve the model serving stack using TensorRT-LLM, Triton Inference Server, vLLM, TGI, continuous batching, KV reuse, speculative decoding, and mixture-of-agents routing.
- Tune distributed training and inference parallelism, including FSDP, ZeRO, tensor parallelism, pipeline parallelism, expert parallelism, and NCCL.
- Implement quantization and PEFT serving with AWQ, GPTQ, FP8, LoRA, and DoRA.
- Build and operate Ray, Kubernetes, and Argo infrastructure with Prometheus, Grafana, and OpenTelemetry observability, autoscaling, A/B infrastructure, canary releases, and rollback mechanisms.
Requirements
- Previous experience at infrastructure-heavy startups such as Databricks or Roblox is highlighted as a technical signal.
- Experience optimizing GPU performance, model serving, distributed parallelism, quantization, observability, autoscaling, and deployment infrastructure is relevant to the scope of work.
Benefits
- On-site, in-person role based in San Mateo.
Tech Stack
Categories
About Moonlake
World models and simulation infrastructure for physical AI.
