21 days ago
New Delhi, IndiaMid Level / Senior
Responsibilities
- Build evaluation infrastructure with audio-native metrics for barge-in, prosody, and turn-taking.
- Create adversarial datasets covering accents and edge cases.
- Develop LLM-as-judge rubrics for task success, tool-use correctness, and recovery.
- Build tracing and observability that correlates audio, STT, LLM reasoning, tool calls, and TTS across a conversation.
- Create analysis and alerting systems that surface cascade failures.
- Mine production traces for failure patterns and generate targeted training or prompt data.
- Validate fixes with adversarial replay and guard against regressions.
- Safeguard sensitive company data and report suspected security incidents according to ISMS policies and procedures.
Requirements
- 3–5 years of experience in ML engineering, research engineering, or applied research.
- Strong Python skills and experience with modern ML tooling.
- Depth in at least two of speech and audio models, LLM agent systems, and evaluation or observability infrastructure.
- Experience shipping a substantial system where research met production.
- Ability to evaluate benchmarks critically and translate research from Interspeech, ACL, or NeurIPS into production systems.
- Real-time systems or telephony experience is preferred.
- Experience with RLHF, DPO, or synthetic data pipelines is preferred.
- Familiarity with enterprise deployment requirements such as SOC 2, PII, and data residency is preferred.
- Publications are welcome but not required.
Benefits
- Senior seat on a small team where research and production are combined.
- Access to enterprise conversation data under proper governance.
- Meaningful equity and autonomy over tooling.
- Support to publish.
- The posting describes the role as a fellowship.
