4 months ago
Milpitas, CA, USASenior
Base Salary
$274k - $304k/yr
Responsibilities
- Own and operate the Ray ecosystem, including KubeRay on GKE, Ray Core scheduling, Plasma, and Ray Data GPU-direct streaming from GCS and S3.
- Configure and operate distributed Ray Train workloads with TorchTrainer, DDP, NCCL, H100 clusters, checkpoint recovery, spot-preemption recovery, and warm-start fine-tuning.
- Build and operate a unified LLM inference mesh using Ray Serve, vLLM, SGLang, and NVIDIA Triton.
- Optimize inference performance through fractional GPU allocation, continuous batching, queue-depth autoscaling, and KV-cache tuning.
- Design capability-, version-, and tenant-based model routing with cost-aware fallback between self-hosted SLMs and cloud LLMs.
- Build Flyte-based reinforcement learning infrastructure using Ray RLlib or custom PPO/GRPO loops with Ray Train.
- Operate model promotion and retraining lifecycles from quality gates and testing through shadow mode, A/B evaluation, canary rollout, and automated rollback.
- Integrate RAG retrieval, vector similarity search, context assembly, and prompt construction into the inference mesh.
Requirements
- Experience in ML engineering, preferably including ML platform or MLOps work.
- Production experience with Ray Train, Ray Serve, Ray Core, and Ray Data, including diagnosing NCCL timeouts, Plasma OOM, and Serve autoscaling issues.
- Hands-on experience with vLLM, SGLang, or NVIDIA Triton, including PagedAttention, prefix caching, and continuous batching.
- Experience with distributed training using DDP, FSDP, NCCL collectives, gradient checkpointing, and BF16 or FP8 mixed precision.
- Working knowledge of reinforcement learning techniques such as PPO, policy gradients, or RLHF.
- Experience with MLflow model registries, shadow/A/B/canary deployments, and golden-signal rollback.
- Experience with Pgvector or Qdrant, ANN indexing, embedding upserts, and query-latency tuning.
- Strong Python and PyTorch skills plus experience with Flyte or an equivalent ML orchestrator.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience.
- Preferred experience with INT8, INT4, or FP8 post-training quantization using GPTQ, AWQ, or bitsandbytes.
Benefits
- Competitive total rewards package.
- Learning and opportunities to grow and advance in the career.
- Potential eligibility for a discretionary bonus plan; compensation decisions consider location, skills, experience, training, licensure, certifications, and business needs.
