Nuna

Software Engineer, AI Evaluation

Nuna
Apply
3 hours ago

Base Salary

$148k - $232k/yr

Responsibilities

  • Build testing harnesses and evaluation infrastructure for agentic products and internal agentic tooling.
  • Own evaluation architecture and content end to end with support from data science and clinical partners.
  • Ensure every agentic deployment passes through the testing apparatus before release and own the release gates that prevent unsafe or low-quality behavior from reaching patients.
  • Build ground truth, judges, and metrics, and validate calibration to human labels, reliability, and confidence in evaluation results.
  • Build labeling and review workflow tools that let clinicians, coaches, and designers author and review evaluation scenarios without engineering support.
  • Connect evaluation results to model and prompt refinement so systems can iterate safely with less manual intervention.

Requirements

  • Significant experience building and shipping reliable production systems and tooling.
  • Deep experience evaluating AI systems in production, including LLM-as-judge, red-teaming, adversarial testing, synthetic scenario generation, and multi-turn or agentic evaluation.
  • A testing mindset focused on building measurement systems, with adversarial instinct, coverage thinking, regression discipline, and clear documentation.
  • Fluency in statistics and experimental design sufficient to partner with a data scientist on calibration and reliability.
  • Ability to design workflows and build functional user interfaces for clinicians, labelers, and other non-engineers.
  • Experience using AI in daily work and building tools that improve team effectiveness.
  • Genuine interest in improving healthcare and judgment to distinguish launch-blocking issues from nice-to-have improvements.
  • Preferred: healthcare or regulated, high-trust domain experience and familiarity with the regulatory landscape.
  • Preferred: hands-on experience with LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar evaluation tooling.
  • Preferred: red-teaming or AI safety experience, including prompt injection, jailbreaks, adversarial testing, and stress testing.
  • Preferred: experience with automated, eval-driven model or prompt optimization.
  • Preferred: experience working in an early-stage or fast-moving environment.
Nuna

About Nuna

51-200 employees

Our data-driven products and services power value-based payment arrangements and facilitate high-value healthcare delivery.