3 months ago
Base Salary
$189k - $283k/yr
Responsibilities
- Own the training lifecycle for production small language models and classifiers, including model selection, supervised fine-tuning, preference optimization, distillation, adversarial training, and post-training optimization.
- Design high-performance inference and serving infrastructure for real-time enforcement and batch workloads processing billions of tokens daily.
- Optimize model deployments using shared GPU pools, KV-cache-aware routing, continuous batching, quantization, speculative decoding, canary releases, shadow traffic, and A/B testing.
- Build synthetic-data pipelines, policy back-testing, online evaluation, drift detection, calibration monitoring, and policy-coverage analysis.
- Create customer-context and memory systems that incorporate data sensitivity, identity, and historical agent behavior into enforcement decisions.
- Mine production agent sessions and customer feedback to identify security gaps, policy improvements, false positives, missed violations, and model failure causes.
- Provide technical leadership for a SAGE model-stack pillar, mentor engineers, shape the roadmap, and collaborate across product, security, platform, research, and customer-facing teams.
Requirements
- Bachelor’s degree or higher in Computer Science, Machine Learning, Computer Engineering, Statistics, or a closely related technical field.
- At least 2 years of professional ML experience with end-to-end production ownership.
- Proficiency in Python and PyTorch or equivalent tools for production training and evaluation.
- Hands-on production experience training, fine-tuning, or distilling language models or classifiers, including supervised fine-tuning and at least one preference-optimization technique such as DPO, RLAIF, or RLHF.
- Production experience with serving frameworks such as vLLM, SGLang, TensorRT-LLM, or equivalent, including continuous batching, KV-cache strategies, and inference-time quantization.
- Experience building closed-loop ML systems covering evaluation, telemetry, data curation, synthetic data, and subsequent model releases.
- Ability to debug and operate high-QPS models in safety-critical request paths.
- Preferred experience in AI safety, red-teaming, adversarial ML, prompt-injection defense, LLM-as-judge evaluation, calibration monitoring, retrieval and context-fusion systems, low-latency inference, weak supervision, active learning, embedding-based retrieval, knowledge distillation, agentic ecosystems, model gateways, and open-source ML contributions.
Categories
About Rubrik
Rubrik builds an enterprise data security and backup platform that protects and recovers data, identities, and workloads across on-prem, public clouds, and SaaS. Its Rubrik Security Cloud delivers backup, ransomware detection and recovery, and policy management, sold primarily by subscription with professional services. Founded in 2014 and headquartered in Palo Alto, California, the company is publicly traded on the NYSE under the ticker RBRK.
