8 months ago
Base Salary
$240k - $280k/yr
Responsibilities
- Design and build evaluation frameworks for accuracy, reliability, regressions, and edge cases in AI systems.
- Create and curate datasets, golden test cases, and benchmarks using real production data.
- Build automated test harnesses and metrics pipelines for continuously evaluating models, prompts, and agentic workflows.
- Partner with applied AI engineers and product leaders to define quality criteria and translate them into measurable tests and metrics.
- Own the evaluation lifecycle for major AI initiatives from early experimentation through production monitoring.
Requirements
- At least 5 years of professional experience and a bachelor’s degree in computer science, machine learning, or a related field.
- Experience building testing, evaluation, or data infrastructure for complex systems, with AI/ML experience strongly preferred.
- Ability to write production-quality code in Python and TypeScript.
- Experience with structured and unstructured datasets, labeling workflows, or data quality pipelines.
- Familiarity with modern ML systems and evaluation techniques, including offline metrics, online evaluation, and regression testing for models or prompts.
- Experience evaluating LLMs, agentic systems, or AI-assisted developer tools is a bonus.
Benefits
- Hybrid work model with Mondays, Tuesdays, and Thursdays as in-office anchor days.
- Eligible for employee benefit plans including paid time off and group health insurance.
- Eligible for incentive compensation and equity grants.
Tech Stack
Categories
About Sentry
Sentry builds application monitoring and error/performance tracking tools for software developers, delivered as a SaaS with open-source SDKs and integrations. Teams use Sentry to detect, triage, and resolve production issues across web and mobile apps; the platform is used by 200,000+ organizations. Founded in 2011 and headquartered in San Francisco, Sentry is a privately held company.
