2 months ago
Bengaluru, IndiaIntern
Responsibilities
- Build evaluation systems for AI features, including failure taxonomies, LLM-as-judge rubrics, golden datasets, and calibration against human judgment.
- Help route tasks across models while balancing cost, quality, and latency.
- Run experiments to justify model-routing decisions and identify regressions.
- Turn noisy real-world signals into reliable, statistically grounded scores and improve heuristic approaches through calibration and monitoring.
- Work on feature pipelines, model versioning, rollouts, drift monitoring, and detection of silent quality degradation.
- Collaborate with engineering on reliable, low-latency model serving and contribute across backend, frontend, data pipelines, and product work.
Requirements
- Currently pursuing or recently completed a degree in computer science, data science, machine learning, or a related field.
- Some hands-on data science or machine learning experience through coursework, personal projects, research, or a prior internship, including building and running something end to end.
- Comfort with Python and working knowledge of SQL.
- Basic applied-statistics knowledge, including the ability to explain what a metric means and when it may be misleading.
- Curiosity about product decisions, backend, or frontend work in addition to modeling.
- Some exposure to LLMs through prompting, APIs, or experimentation with model behavior.
- Preferred: exposure to evaluation or observability tooling for LLM features.
- Preferred: coursework or projects in information retrieval, entity matching, or record linkage.
- Preferred: interest in developer productivity, code analytics, or DevEx data.
