3 months ago
Remote, AmericasSenior
Responsibilities
- Build visual-data pipelines that extract structure, layout, charts, and context from complex documents.
- Use vision-language models and document-intelligence tools to interpret visual documents.
- Develop agents and chains with LangChain or LangGraph for querying, reasoning, and response generation.
- Process PDFs natively while preserving document structure before AI processing.
- Create prompts and control flows that improve model accuracy and reduce hallucinations.
- Optimize batch processing, model selection, and orchestration costs for scale.
- Collaborate with teammates while contributing to an inclusive and respectful engineering culture.
Requirements
- Deep experience with LangChain, LangGraph, or similar LLM orchestration frameworks, including context windows, tool calling, and agentic workflows.
- Hands-on experience integrating vision models such as GPT-4V and Claude 3.5 Sonnet, plus embedding models such as CLIP.
- Familiarity with Donut, Pix2Struct, Unstructured.io, or Docling for document intelligence.
- Strong experience with PDF processing tools such as PyMuPDF or pdfplumber.
- Strong proficiency in Python ML frameworks such as PyTorch or TensorFlow.
- Strong verbal and written communication skills in English.
- Experience fine-tuning vision or language models for domain-specific artifacts such as financial charts or tables is a plus.
- Experience with Real Estate or Finance documents is a plus.
Benefits
- 100% remote within LatAm.
- 40-hour work week with availability during normal business hours as needed.
- Payments made in USD.
- 18 days of PTO per year, local holiday observance, and an annual break between Christmas and New Year’s.
- Monthly wellness stipend and snack boxes delivered to the employee’s home.
- Access to AI training, knowledge-sharing, and hands-on experimentation.
