Moonlake

Member of Technical Staff - Data & ML Infra Engineer

Moonlake
Apply
1 year ago

Responsibilities

  • Optimize CUDA and Triton GPU kernels, FlashAttention, paged attention, and CUDA Graphs.
  • Improve the model serving stack using TensorRT-LLM, Triton Inference Server, vLLM, TGI, continuous batching, KV reuse, speculative decoding, and mixture-of-agents routing.
  • Tune distributed training and inference parallelism, including FSDP, ZeRO, tensor parallelism, pipeline parallelism, expert parallelism, and NCCL.
  • Implement quantization and PEFT serving with AWQ, GPTQ, FP8, LoRA, and DoRA.
  • Build and operate Ray, Kubernetes, and Argo infrastructure with Prometheus, Grafana, and OpenTelemetry observability, autoscaling, A/B infrastructure, canary releases, and rollback mechanisms.

Requirements

  • Previous experience at infrastructure-heavy startups such as Databricks or Roblox is highlighted as a technical signal.
  • Experience optimizing GPU performance, model serving, distributed parallelism, quantization, observability, autoscaling, and deployment infrastructure is relevant to the scope of work.

Benefits

  • On-site, in-person role based in San Mateo.

Tech Stack

Argo CDGrafanaKubernetesPrometheus
Moonlake

About Moonlake

11-50 employees

Moonlake builds interactive world models and simulation infrastructure that generate, simulate, and reason over 3D environments for embodied AI, robotics, and gaming. Its platform and APIs let researchers and developers create assets, scenes, and digital twins at scale, and interact with them via natural language and multimodal inputs. The company is privately held, founded in 2025, headquartered in San Francisco, and raised a $28M seed round in 2025.

Contact me