3 months ago
Base Salary
$300k - $400k/yr
Responsibilities
- Set up end-to-end evaluations to measure and improve agent performance.
- Experiment with multi-agent systems, reasoning-from-feedback, and other agentic techniques.
- Build lightweight tools, servers, and orchestration layers, including MCP servers, for reliable production agents.
- Research emerging LLM and agent techniques and bring relevant ideas into production experiments.
- Develop evaluation harnesses, orchestration and tools, multi-agent systems, reasoning loops, and long-running agents that work with documents and connected systems.
Requirements
- Strong experience with Python.
- At least 8 years of software engineering experience, including at least 3 years building in AI/ML.
- Ability to work effectively with LLMs through clear, concise prompting.
- Clear written and verbal communication skills.
- Bonus experience in B2B SaaS and 0-to-1 product development.
- High drive, grit, ownership, and ability to work in a fast-paced, high-expectation environment.
Benefits
- Fully covered health, dental, and vision benefits.
- Competitive compensation, meaningful stock options, and a company 401(k).
- Unlimited paid time off.
- Outstanding in-office culture in San Francisco.
- Lunch and dinner provided onsite.
- Team events including happy hours and off-sites.
- Pre-tax commuter benefits.