5 months ago
Bengaluru, IndiaSenior
Responsibilities
- Design and build agentic evaluation pipelines covering error detection, root-cause analysis, hypothesis generation, prompt testing, A/B measurement, and production promotion.
- Own accuracy measurement infrastructure, including error analysis, data quality pipelines, batch evaluation frameworks, classification and extraction correction loops, and performance reporting.
- Build and productionize LLM-based extraction and classification pipelines using few-shot prompting and RAG strategies.
- Design A/B testing infrastructure and live dashboards for accuracy, NTP rates, and false positives across document types and customer configurations.
- Optimize LLM costs through prompt compression, output-token minimization, and model selection or migration strategies.
- Develop reliable production data pipelines with error handling, retries, logging, and monitoring.
- Collaborate with platform engineering and applied research teams and mentor one to two junior engineers.
Requirements
- Bachelor’s or master’s degree in Computer Science, AI/ML, Computational Data Science, Computer Science & Automation, or a related discipline.
- 8–10 years of total experience, including 4–6 years building production LLM or AI systems and 4–6 years in evaluation, quality measurement, or accuracy improvement.
- Production-grade Python experience, including pytest and FastAPI or Flask.
- Hands-on experience with LLM APIs such as OpenAI, Anthropic, Gemini, or AWS Bedrock and systematic prompt engineering.
- Experience designing agentic pipelines with tool use, orchestration frameworks, automated evaluation, and feedback loops.
- Knowledge of LLM evaluation metrics including precision, recall, F1, confusion matrices, A/B testing, and per-class error analysis.
- Experience with MongoDB or equivalent NoSQL systems, including queries, aggregations, and indexing.
- Experience with pandas or NumPy for data processing and batch analysis, plus Git, code reviews, and CI/CD basics using GitHub Actions or Jenkins.
- Preferred experience with document AI, PDF parsing, layout-aware extraction, OCR, structured form extraction, RAG and vector search, large-label-space classification, asynchronous Python, embeddings, and semantic similarity.
- Experience working alongside a research or applied science team as an engineering counterpart is preferred.
