about 4 hours ago
Responsibilities
- Build extraction pipelines that transform documents, transcripts, and records into structured, schema-validated data.
- Combine LLMs with traditional parsing techniques, including OCR and layout analysis, for reliable extraction.
- Design validation rules and confidence-scoring systems for low-confidence cases.
- Build evaluation datasets and testing frameworks to measure extraction accuracy.
- Process diverse inputs, including clean PDFs and poor-quality scans.
- Collaborate with clients and product teams to ensure clean data flows.
- Communicate pipeline capabilities and risks to stakeholders.
- Use AI coding agents as part of daily engineering workflows.
Requirements
- 5+ years of experience shipping production software.
- Hands-on experience extracting structured data from unstructured sources using LLMs, NLP, or OCR.
- Experience with LLM structured outputs and prompt iteration.
- Proficiency in TypeScript and Node.js, or strong experience in Python.
- Familiarity with document AI tools like Amazon Textract or Google Document AI.
- Experience building solutions on AWS.
- Strong understanding of edge cases and confidence scoring.
- Professional English proficiency at B2 level or higher.
- Strong written and verbal communication skills.
- Experience in fast-paced, agile consulting environments.
Tech Stack
Categories
AI & MLData Engineering
