3 hours ago
Base Salary
$148k - $232k/yr
Responsibilities
- Build testing harnesses and evaluation infrastructure for agentic products and internal agentic tooling.
- Own evaluation architecture and content end to end with support from data science and clinical partners.
- Ensure every agentic deployment passes through the testing apparatus before release and own the release gates that prevent unsafe or low-quality behavior from reaching patients.
- Build ground truth, judges, and metrics, and validate calibration to human labels, reliability, and confidence in evaluation results.
- Build labeling and review workflow tools that let clinicians, coaches, and designers author and review evaluation scenarios without engineering support.
- Connect evaluation results to model and prompt refinement so systems can iterate safely with less manual intervention.
Requirements
- Significant experience building and shipping reliable production systems and tooling.
- Deep experience evaluating AI systems in production, including LLM-as-judge, red-teaming, adversarial testing, synthetic scenario generation, and multi-turn or agentic evaluation.
- A testing mindset focused on building measurement systems, with adversarial instinct, coverage thinking, regression discipline, and clear documentation.
- Fluency in statistics and experimental design sufficient to partner with a data scientist on calibration and reliability.
- Ability to design workflows and build functional user interfaces for clinicians, labelers, and other non-engineers.
- Experience using AI in daily work and building tools that improve team effectiveness.
- Genuine interest in improving healthcare and judgment to distinguish launch-blocking issues from nice-to-have improvements.
- Preferred: healthcare or regulated, high-trust domain experience and familiarity with the regulatory landscape.
- Preferred: hands-on experience with LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar evaluation tooling.
- Preferred: red-teaming or AI safety experience, including prompt injection, jailbreaks, adversarial testing, and stress testing.
- Preferred: experience with automated, eval-driven model or prompt optimization.
- Preferred: experience working in an early-stage or fast-moving environment.
Categories
About Nuna
Our data-driven products and services power value-based payment arrangements and facilitate high-value healthcare delivery.
