21 days ago
New Delhi, IndiaEntry Level
Responsibilities
- Design audio-native evaluation metrics for barge-in, prosody, latency-induced errors, and cross-turn context loss.
- Generate adversarial conversational datasets across accents and edge cases.
- Build LLM-as-judge rubrics for task completion, empathy, and recovery from tool failures.
- Shape schemas and analysis layers that correlate audio, speech recognition, LLM, tool-call, and text-to-speech traces.
- Mine production traces for failure patterns and generate targeted fine-tuning data or prompt updates.
- Validate agent improvements through adversarial replay.
- Contribute to the intelligence layer for self-healing enterprise voice agents.
Requirements
- Comfortable programming in Python.
- Familiarity with at least one of speech models such as Whisper or Conformer variants, LLM tool-use and agent frameworks, or observability stacks such as OpenTelemetry, Langfuse, Arize, or Hamming.
- Current PhD study in machine learning, NLP, or speech is the strong default; exceptional MS students or research engineers with a publication track record are welcome.
- Relevant research or publications in speech, dialogue systems, agent evaluation, or human-AI interaction are valued.
- Prior work on evaluation methodology, dataset synthesis, or interpretability is preferred.
- Experience with real-time systems, telephony, or streaming pipelines is preferred.
- A blog, repository, or workshop paper demonstrating technical thinking is preferred.
Benefits
- Full-time role with competitive salary and ESOPs.
- Fun, flexible, and creative work environment with high ownership.
- Play area, free lunches, and complimentary chai and coffee.
- Workspace based in ixigo's entrepreneurial startup environment.
