Base Salary
$137k - $202k/yr
Responsibilities
- Lead the design of infrastructure that moves generative AI ideas from prototype to production.
- Own and evolve real-time GPU endpoints, high-throughput batch inference, fine-tuning pipelines, the LLM Gateway, Agent Gateway, evaluation infrastructure, guardrails, and cost attribution.
- Architect scalable systems for model serving, batch inference, GPU autoscaling, fine-tuning, backend services, and observability.
- Improve GPU inference cost, latency, throughput, batching, autoscaling, and utilization while supporting open-weight and closed-source model choices.
- Build production-grade platforms with monitoring, SLOs, operational playbooks, reliability, fallback, and cost controls.
- Partner with ML engineers, product engineers, data scientists, and platform teams across DoorDash, Wolt, and Deliveroo.
- Set technical direction for future generative AI capabilities, including reinforcement learning, agent optimization, post-training, and agentic techniques.
- Mentor engineers and raise the technical bar across the team.
Requirements
- B.S., M.S., or PhD in Computer Science or equivalent experience.
- At least 6 years of industry software engineering experience.
- Deep backend engineering fundamentals, especially Python and distributed systems.
- Experience designing and owning production services, APIs, data pipelines, or ML infrastructure at scale.
- Experience operating production systems, including observability, debugging, reliability, incident response, and performance and cost optimization.
- Hands-on production experience with LLM inference and/or fine-tuning open-weight models, including serving, batching, autoscaling, GPU utilization, SFT, DPO, or LoRA.
- Demonstrated technical leadership across ambiguous technical areas, including design leadership, mentoring, and creating reusable platform capabilities.
- Proficiency using AI coding tools such as Claude Code, Codex, or Cursor across the software development lifecycle.
- Experience with inference engines and serving frameworks such as vLLM, SGLang, or TensorRT-LLM is preferred.
- Experience with distributed or multi-node fine-tuning and training pipelines, GPU performance optimization, Kubernetes, AWS, GCP, Modal, LLM gateways, developer platforms, AI agents, MCP servers, evaluation systems, observability, tracing, RAG, search, or vector databases is preferred.
Benefits
- 401(k) plan with employer matching
- 16 weeks of paid parental leave
- Wellness benefits and expense reimbursement
- Commuter benefits match
- Flexible paid time off/vacation for salaried roles
- Paid sick leave
- Medical, dental, and vision benefits
- 11 paid holidays
- Disability and basic life insurance
- Family-forming assistance and mental health program
- Equity grant opportunities
Tech Stack
Categories
About DoorDash
At DoorDash, our mission to empower local economies shapes how our team members move quickly and always learn and reiterate to support merchants, Dashers and the communities we serve. We are a technology and logistics company that started with door-to-door delivery, and we are looking for team members who can help us go from a company that is known for delivering food to a company that people turn to for any and all goods. DoorDash is growing rapidly and changing constantly, which gives our team members the opportunity to share their unique perspectives, solve new challenges, and own their careers. Our leaders seek the truth and welcome big, hairy, audacious questions. We are grounded in our company values, and we make intentional decisions that are both logical and display empathy for our range of users—from Dashers to Merchants to Customers.