22 hours ago
Amsterdam, NetherlandsSenior
Base Salary
$225k - $255k/yr
Responsibilities
- Build evaluation harnesses and benchmarks using tracked pricing outcomes as ground truth for model quality.
- Automate expert review workflows and persona training pipelines.
- Develop AI personas that simulate B2B buying committees using usage data and call transcripts.
- Own LLM routing across Anthropic, Google, and other providers while balancing cost, latency, and quality.
- Maintain infrastructure and data residency boundaries, including EU-only model calls where required.
- Extend the MCP server used by LLM agents so product features are agent-driven.
- Use typed pricing ontologies to keep model outputs structured and auditable.
- Identify and remediate systemic latency, data drift, and cold-start issues.
- Ship and own LLM-powered product features from development through production monitoring.
Requirements
- 8 or more years of engineering experience with strong, recent production-grade LLM experience.
- Proven experience shipping and owning LLM-powered product features end-to-end.
- Hands-on experience building LLM evaluation harnesses and observability systems.
- Experience routing across multiple LLM providers with explicit cost, latency, and quality tradeoffs.
- Experience with structured data models and typed ontologies, such as Pydantic models.
- Familiarity with MCP or building tools for LLM agents.
- Familiarity with LangChain, LlamaIndex, Braintrust, or OpenRouter.
- Experience operating under data residency, SOC2, and GDPR constraints.
- Pricing, billing, or payments domain experience is preferred.
- Ability to communicate clearly about non-deterministic systems to clients and non-technical stakeholders.
- Must be authorized to work in the United States.
Benefits
- On-site role in Amsterdam, North Holland, Netherlands.
- Candidates must be based in Amsterdam or willing to relocate.
