29 days ago
Bengaluru, IndiaMid Level
Responsibilities
- Design, build, test, and iterate on prompts and prompt chains for production LLM workflows.
- Implement evaluation harnesses, guardrails, and regression tests to measure and stabilize output quality.
- Integrate LLM components with business systems and data through APIs, retrieval/RAG pipelines, and structured outputs.
- Partner with AI Specialists to convert solution designs into production builds and document reusable patterns.
- Tune solutions for accuracy, cost, latency, and safety using real evaluation data.
- Develop Caravel’s library of reusable AI accelerators.
Requirements
- At least 2 years of experience building software or AI/LLM components, including hands-on prompt development for real applications.
- Working knowledge of context windows, few-shot prompting, function/tool calling, structured outputs, and RAG basics.
- Proficiency in Python or a similar language and familiarity with APIs, JSON, and Git.
- A test-and-measure approach to evaluating output quality.
- Clear written English and strong documentation habits.
- Ability to collaborate across time zones with an onshore/nearshore team.
- Preferred experience with promptfoo, Ragas, or custom evaluation harnesses; embeddings, vector databases, and retrieval pipelines; NetSuite, Salesforce, or business-process automation; and agentic frameworks and orchestration.
