3 hours ago
Base Salary
$160k - $250k/yr
Responsibilities
- Own agent task-success metrics, establish performance baselines, and improve workflow completion rates.
- Build reproducible evaluation infrastructure grounded in validated user stories and real engineering workflows.
- Set per-task token budgets and track the cost of completed workflows.
- Map, interview about, and validate user workflows with researchers and engineering domain experts.
- Translate user stories into testable evaluations and prioritize workflow coverage.
- Make architecture decisions across tool calling, state management, error recovery, model routing, and context management.
- Set technical direction for a small team, review designs and code, unblock teammates, and contribute production code.
- Collaborate with product, integrations, and customers to align agent behavior with real-world use.
Requirements
- At least 7 years of software engineering experience, including 2 or more years building and shipping LLM-based agents that take real-world actions.
- Deep experience with LLM application architecture, including model selection, context management, retrieval, tool calling, and orchestration.
- Strong Python skills and familiarity with function calling, tool APIs, tracing or observability, and evaluation tooling.
- Experience building benchmarks for task completion, cost efficiency, and failure analysis.
- Hands-on technical leadership and code review experience on a small engineering team.
- Experience applying AI or LLM tooling to proprietary engineering data or desktop engineering software.
- Background in mechanical engineering, CAD, CAE, PLM, or a related domain.
- Familiarity with desktop automation and enterprise workstation constraints is valuable.
Benefits
- Salary range of $160,000 to $250,000 USD annually.
- On-site work in San Francisco, California.
- Visa sponsorship is not available.
