3 months ago
Responsibilities
- Design, build, ship, and operate production agentic systems and workflows end to end.
- Automate high-leverage SME workflows such as review pipelines, content QA, labeling, support triage, and internal copilots.
- Build offline evaluations, calibrated LLM-as-judge systems, regression suites, online metrics, and versioned golden datasets.
- Design safety, bias, privacy, compliance, human-in-the-loop, and content-filtering controls for user-facing AI features.
- Hand off shipped features through documentation, runbooks, evaluation harnesses, dashboards, and receiving-team enablement.
- Develop internal AI libraries, patterns, and playbooks that help other engineering teams ship AI features independently.
Requirements
- 7+ years of professional engineering experience, including at least 1 year building and operating agentic systems in production.
- Hands-on experience with an agentic stack such as LangGraph, Anthropic SDK/Claude Agent SDK, OpenAI Agents SDK, Pydantic-AI, Mastra, LlamaIndex, CrewAI, or a comparable homegrown stack.
- Strong Python skills plus a typed production language such as TypeScript or Go, along with AWS or GCP and containerized deployment experience.
- Senior- or staff-level software engineering foundation and a track record of leading production systems to launch.
- Multiple shipped LLM-powered production features with experience diagnosing failures and improving systems.
- Practical knowledge of agentic patterns, retrieval, grounding, prompt engineering, structured output validation, and disciplined evaluation practices.
- Strong written and verbal communication skills and comfort using AI coding tools while owning system design, verification, and delivery.
Benefits
- Onsite five days per week in Austin, Texas.
- Relocation assistance is offered.
- Typical embedded engagements last 4–12 weeks.
- The interview process includes an AI-assisted coding session using the candidate’s own development setup and AI tools.
