Responsibilities
- Design, build, and own infrastructure for distributed ML training and low-latency, highly available inference.
- Implement and scale LLM inference stacks while optimizing throughput, latency, token streaming, GPU utilization, and automated scaling.
- Partner with AI Research and Data Science teams to accelerate model experimentation, fine-tuning, deployment, and developer productivity.
- Build automated workflows for model validation, deployment, continuous training, and lifecycle management.
- Develop observable and resilient ML systems using modern monitoring and telemetry tools.
Requirements
- 5+ years of infrastructure or software engineering experience, including at least 2+ years focused on MLOps or ML infrastructure for large-scale distributed systems.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
- Deep hands-on production experience with Kubernetes and cloud-native tools including Helm, ArgoCD, and Argo Workflows.
- Experience optimizing GPU resource utilization, data ingestion, model training, and deployment scalability.
- Hands-on experience with LLM inference serving frameworks such as vLLM, SGLang, Triton Inference Server, or Ray Serve.
- Strong programming proficiency in Python or Go and experience with PyTorch, Jax, or TensorFlow.
- Experience with observability and resilient systems using tools such as Prometheus, Grafana, or OpenTelemetry.
- Preferred: experience writing custom inference kernels in CUDA or Triton, applying quantization, distillation, or speculative decoding, serving multimodal models, and using AI safety or evaluation frameworks.
Tech Stack
Categories
About DevRev
Computer is the most precise, efficient, and safe AI platform for enterprise search, employee service/help desk, and customer support (including voice). It works by unifying all your existing systems and data into a patented, AI-ready knowledge graph: Native Shared Memory. Every tool, every team, every customer conversation lives in this compounding context. So every answer and action is grounded in what your company knows, and how it operates. With permissions and safety guardrails respected, always. PROVEN EFFICIENCY, PRECISION, AND SAFETY We tested Computer against the market-leading AI on enterprise data: – 48% more accurate – 4.4x fewer tokens used – Cost stays flat at 256x data growth (vs. 29% growth) These results are from “Enterprise-Bench” – the AI industry’s first open, vendor-neutral benchmarking tool. Developed with Laude Institute, validated by Professor Alexandros Dimakis (UC Berkeley). The benchmark model, dataset, and traces are all public. (Feel free to get in touch to learn more.) REAL-WORLD RESULTS More and more companies are trusting Computer: – Bolt: 40% faster ticket resolution. 43% cut in live ticket volume. 25% increase in customer retention. – Bill: 70% of support tickets fully resolved without a human. 100% deployment across all market segments in 15 weeks. Up to $6M in annual savings. – $1.1B org*: 2,000+ new hires onboarded in one year without hiring extra IT staff. 10 weeks to full deployment across all teams. *We can’t reveal the name of our (very) happy customer just yet – but we will soon. DevRev is headquartered in Palo Alto, with eight global offices. We’re backed by Khosla Ventures and Mayfield, with $150M+ raised. Founded by Dheeraj Pandey, former co-founder and CEO of Nutanix [NASDAQ: NTNX], and board member at Adobe [NASDAQ: ADBE], alongside Manoj Agarwal, former SVP of Engineering at Nutanix.