Saviynt

AI Platform Engineer, Training and Inference

Saviynt
Apply
4 months ago
Milpitas, CA, USASenior

Base Salary

$274k - $304k/yr

Responsibilities

  • Own and operate the Ray ecosystem, including KubeRay on GKE, Ray Core scheduling, Plasma, and Ray Data GPU-direct streaming from GCS and S3.
  • Configure and operate distributed Ray Train workloads with TorchTrainer, DDP, NCCL, H100 clusters, checkpoint recovery, spot-preemption recovery, and warm-start fine-tuning.
  • Build and operate a unified LLM inference mesh using Ray Serve, vLLM, SGLang, and NVIDIA Triton.
  • Optimize inference performance through fractional GPU allocation, continuous batching, queue-depth autoscaling, and KV-cache tuning.
  • Design capability-, version-, and tenant-based model routing with cost-aware fallback between self-hosted SLMs and cloud LLMs.
  • Build Flyte-based reinforcement learning infrastructure using Ray RLlib or custom PPO/GRPO loops with Ray Train.
  • Operate model promotion and retraining lifecycles from quality gates and testing through shadow mode, A/B evaluation, canary rollout, and automated rollback.
  • Integrate RAG retrieval, vector similarity search, context assembly, and prompt construction into the inference mesh.

Requirements

  • Experience in ML engineering, preferably including ML platform or MLOps work.
  • Production experience with Ray Train, Ray Serve, Ray Core, and Ray Data, including diagnosing NCCL timeouts, Plasma OOM, and Serve autoscaling issues.
  • Hands-on experience with vLLM, SGLang, or NVIDIA Triton, including PagedAttention, prefix caching, and continuous batching.
  • Experience with distributed training using DDP, FSDP, NCCL collectives, gradient checkpointing, and BF16 or FP8 mixed precision.
  • Working knowledge of reinforcement learning techniques such as PPO, policy gradients, or RLHF.
  • Experience with MLflow model registries, shadow/A/B/canary deployments, and golden-signal rollback.
  • Experience with Pgvector or Qdrant, ANN indexing, embedding upserts, and query-latency tuning.
  • Strong Python and PyTorch skills plus experience with Flyte or an equivalent ML orchestrator.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical or military experience.
  • Preferred experience with INT8, INT4, or FP8 post-training quantization using GPTQ, AWQ, or bitsandbytes.

Benefits

  • Competitive total rewards package.
  • Learning and opportunities to grow and advance in the career.
  • Potential eligibility for a discretionary bonus plan; compensation decisions consider location, skills, experience, training, licensure, certifications, and business needs.

Tech Stack

Google CloudKubernetesMLflowPythonPyTorch
Saviynt

About Saviynt

1,001-5,000 employees
Contact me