3 months ago
Base Salary
$160k - $250k/yr
Responsibilities
- Lead development of the core agent intelligence layer for multi-step workflows across CAD, CAE, and PLM software.
- Serve as technical lead for a small team of AI engineers, a user researcher, and domain expert contractors.
- Own the product loop from user-story definition through implementation and benchmarking against real workflows.
- Define evaluation frameworks, establish baselines, improve task completion metrics, and track token budgets and workflow costs.
- Build reproducible evaluation infrastructure grounded in validated user stories.
- Lead user-story mapping and validation through interviews and collaboration with domain experts.
- Own agent architecture decisions covering tool calling, state management, error recovery, model routing, and context management.
- Write production code, review designs, unblock the team, and raise engineering standards.
- Collaborate with integrations, product, and customers during proof-of-concept engagements.
Requirements
- 7+ years of software engineering experience, including at least 2 years building agentic LLM-based agents for real-world, multi-step workflows.
- Deep experience designing LLM application architectures, including model selection, context-window management, retrieval strategies, tool-calling frameworks, and orchestration patterns.
- Strong evaluation and benchmarking experience for agentic systems, including task completion, cost efficiency, and failure-mode analysis.
- Proven experience shipping AI systems with measurable outcomes rather than demonstrations alone.
- Strong Python skills and hands-on experience with LLM tooling, function calling, tool-use APIs, tracing and observability tools such as Logfire or LangSmith, and evaluation frameworks.
- Experience leading a small technical team of 3–6 engineers, setting direction, conducting code reviews, and driving architecture decisions.
- Preferred: experience with desktop automation, COM, or programmatic application control beyond web APIs.
- Preferred: background in mechanical engineering, CAD/CAE, PLM, or adjacent industries.
- Preferred: familiarity with enterprise deployment constraints on locked-down corporate workstations.
- Preferred: published work or open-source contributions in agentic AI systems.
- Preferred: experience building or contributing to public benchmarks for AI agents.
Benefits
- Equity participation in an early-stage Series A company.
- On-site role in San Francisco, California; not remote.
