12 hours ago
Palo Alto, CA, USASenior
Base Salary
$200k - $300k/yr
Responsibilities
- Build autonomous agents that complete delegated executive tasks end to end without human follow-up.
- Design and maintain evaluation frameworks, success metrics, datasets, and readiness evidence for production agent launches.
- Expand the tasks agents can handle in tools used by executives.
- Turn production failures into permanent fixes and build detection for known and previously unseen failure patterns.
- Build self-recovery capabilities for failures involving tools or external services.
- Run realistic end-to-end simulations and measure agent improvements.
Requirements
- At least 5 years of experience in data science, machine learning, or analytics focused on evaluation systems and quality measurement for production AI.
- Experience designing evaluation methodologies, including success criteria, dataset construction, metric selection, and signal-versus-noise analysis.
- Production-quality Python and SQL skills, with familiarity with Django, React, and TypeScript.
- Strong statistical and experimental design skills, including sampling, variance, bias detection, and significance testing for non-deterministic systems.
- Experience with LLM-as-a-judge systems, model-based graders, and grader calibration.
- Ability to debug across prompts, traces, model outputs, code, databases, and APIs.
- Experience building and operating production agentic systems with multi-step execution and tool use.
- Experience developing ground-truth data, labeling guidelines, annotation quality control, and dataset maintenance.
- Strong computer science or engineering fundamentals from a rigorous degree program.
- Demonstrated ownership of projects shipped to real users and measurable impact from those projects.
Benefits
- On-site work in Palo Alto, California.
- Salary range of $200,000 to $300,000 USD annually.
