1 month ago
Buenos Aires, ArgentinaMid Level
Base Salary
$30k - $42k/yr
Responsibilities
- Design, build, test, and iterate on prompts and prompt chains for production LLM workflows.
- Implement evaluation harnesses, guardrails, and regression tests to measure and stabilize output quality.
- Integrate LLM components with business systems and data through APIs, retrieval/RAG pipelines, and structured outputs.
- Partner with AI Specialists to convert solution designs into production builds and document reusable patterns.
- Tune AI systems for accuracy, cost, latency, and safety using real evaluation data.
- Grow the library of reusable AI accelerators.
Requirements
- At least 2 years of experience building software or AI/LLM components, including hands-on prompt development for real applications.
- Working knowledge of context windows, few-shot prompting, function/tool calling, structured outputs, and RAG basics.
- Proficiency in Python or a similar language and comfort with APIs, JSON, and Git.
- A test-and-measure mindset focused on evaluating output quality.
- Clear written English, strong documentation habits, and the ability to collaborate across time zones.
- Preferred experience with promptfoo, Ragas, custom evaluation harnesses, embeddings, vector databases, retrieval pipelines, NetSuite, Salesforce, business-process automation, agentic frameworks, or orchestration.
Benefits
- Fully remote within South America with daily overlap with US business hours.
- Competitive base compensation commensurate with experience.
- Opportunity to help build a new AI practice and reusable AI accelerators.
