2 days ago
Pune, IndiaStaff+
Responsibilities
- Architect and lead cross-product GenAI platform capabilities including LLM Proxy, model registry integrations, vendor abstraction, and cost and usage attribution.
- Design and scale evaluation and benchmarking frameworks, including offline, continuous regression, and A/B evaluation, to gate model releases.
- Define and drive adoption of company-wide standards, rubrics, and automated checks for AI safety, tone, reasoning, quality, and trust.
- Identify systemic failure modes and prioritize mitigations, monitoring, and retraining strategies with ML teams.
- Drive reliability, observability, capacity planning, rate limiting, throttling, and SLA practices for LLM services.
- Enable agentic workflows and safe tool use by defining integration patterns and security boundaries.
- Translate risk, cost, and quality tradeoffs into platform decisions with engineering, product, research, and legal/policy stakeholders.
- Mentor senior engineers, coordinate cross-team roadmaps, and represent the platform in technical forums.
Requirements
- 8+ years of experience building distributed systems and ML infrastructure, including major production responsibilities and large cross-team project delivery.
- Deep understanding of LLMs, inference serving patterns, vendor routing strategies, and ML platform design.
- Strong system design, service reliability engineering, capacity planning, and cost optimization skills.
- Proficiency in Python or a comparable server-side language, Kubernetes, cloud infrastructure, and observability tooling.
- Experience creating evaluation frameworks, gold-standard datasets, regression suites, and monitoring for language or ML systems.
- Demonstrated experience with cloud-native infrastructure such as Kubernetes, AWS, GCP, or Azure and production ML/LLM systems.
- Proven ability to lead technical strategy, influence broad adoption, and mentor senior engineers.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Preferred qualifications include experience with model registries, feature stores, inference platforms, agentic AI frameworks, workflow orchestration, tool-using models, and company-wide ML safety or quality frameworks.
- An advanced degree in ML/NLP or a related field and/or relevant published research is preferred.
Benefits
- Hybrid work with part of the week in the local office; the specific in-office schedule is determined by the hiring manager.
- The role is based in Pune, India, and candidates must be physically located and plan to work from Karnataka or Maharashtra.
- Zendesk offers an inclusive workplace, equal employment opportunity, diversity and inclusion initiatives, and reasonable accommodations for applicants with disabilities.
Categories
About Zendesk
Move beyond deflection. Deliver real resolutions with self-improving AI agents that learn, adapt, and outperform, on every channel, on any platform.
