DigitalOcean

Senior Forward Deployed Engineer I (AI/ML)

DigitalOcean
Apply
3 months ago
Bengaluru, IndiaSenior

Responsibilities

  • Architect, deploy, optimize, and scale production AI and agentic systems for strategic customers and AI startups on DigitalOcean’s AI-Native Cloud.
  • Support complex migrations, production-ready proofs of concept, deployment acceleration, and long-term workload expansion across inference and runtime platforms.
  • Optimize distributed inference and runtime performance through benchmarking, GPU efficiency tuning, KV-cache optimization, speculative decoding, prefill/decode disaggregation, and multi-node deployments.
  • Validate AI-native platform capabilities as a first customer and communicate operational insights, architectural gaps, and scaling bottlenecks to Product Engineering and Research teams.
  • Build deployment frameworks, benchmarking systems, automation tooling, AI starter kits, fine-tuning workflows, operational playbooks, and reference architectures.
  • Collaborate with GPU vendors, model providers, infrastructure partners, and ISVs on co-development, technical validation, optimization, and launch readiness.
  • Enable customer-facing technical and partner teams through deployment patterns, benchmarking insights, playbooks, reference architectures, demos, and technical guidance.
  • Travel for customer engagements, strategic workshops, conferences, and internal collaboration as needed.

Requirements

  • Experience designing and operationalizing production AI systems, including inference workloads, agentic runtimes, orchestration frameworks, and AI-native applications.
  • Hands-on experience with inference and serving frameworks such as vLLM, SGLang, Ray Serve, NVIDIA Dynamo, or llm-d, plus LLM optimization techniques including continuous batching, quantization, KV-cache optimization, and speculative decoding.
  • Deep expertise with NVIDIA and AMD GPU platforms and ecosystems including CUDA, ROCm, TensorRT, Triton, NCCL, RCCL, NVLink, XGMI, and RoCE.
  • Strong proficiency with Kubernetes, distributed systems, networking, storage systems, Infrastructure as Code, and large-scale AI infrastructure architectures.
  • Experience with AI orchestration and agent frameworks such as LangGraph, CrewAI, MCP ecosystems, LlamaIndex, and OpenAI Agents SDK.
  • Strong production coding skills in Python or Go, including tooling, automation systems, deployment workflows, benchmarking frameworks, and operational platforms.
  • Ability to benchmark and optimize AI infrastructure for scalability, reliability, GPU efficiency, runtime performance, latency, and workload economics.
  • Ability to establish technical credibility with CTOs, principal architects, product engineering teams, and ecosystem partners while managing high-impact deployments and initiatives.
  • 4+ years of experience in Forward Deployed Engineering, ML Engineering, Applied AI Engineering, AI Infrastructure, Technical Consulting, or equivalent customer-facing engineering roles supporting production AI systems.
  • Experience building deployment standards, technical enablement programs, platform adoption frameworks, or ecosystem integration strategies.
  • Active contribution to open-source AI, infrastructure, orchestration, or developer tooling ecosystems is preferred.
  • Experience collaborating with GPU vendors, infrastructure providers, model vendors, or ecosystem partners on benchmarking, optimization, technical validation, or launch readiness is preferred.

Benefits

  • Hybrid role located in Bengaluru, India.
  • Travel up to 30% is required, with consistent overlap with North American business hours through at least noon Eastern Time.
  • Reimbursement for relevant conferences, training, and education.
  • Access to LinkedIn Learning courses.
  • Employee Assistance Program, local employee meetups, and flexible time off.
  • Eligible employees may receive equity compensation and participate in the Employee Stock Purchase Program.

Categories

Forward Deployed
DigitalOcean

About DigitalOcean

1,001-5,000 employees

DigitalOcean provides cloud infrastructure and platform services for developers, startups, and small to mid-sized businesses, including virtual machines (Droplets), managed Kubernetes and databases, object/block storage, networking, and GPUs for AI workloads. It operates a usage-based, self-service public cloud with APIs, CLI, and a marketplace to deploy and scale applications. Founded in 2012 and headquartered in Broomfield, Colorado, DigitalOcean is a public company listed on the NYSE.

Contact me