
Sr. Evaluation Engineer
LogicMonitor17 days ago
Base Salary
$158k - $218k/yr
Responsibilities
- Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.
- Build Python-based offline and online evaluation pipelines integrated with experimentation, model selection, prompt iteration, CI/CD, and release gating.
- Create and maintain golden datasets, regression suites, customer-specific scenarios, adversarial test cases, and evaluation-driven safeguards.
- Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows, including planning, retrieval, evidence use, tool selection, and escalation decisions.
- Evaluate AI capabilities such as incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, and automated resolution.
- Design and calibrate LLM-based graders against expert human judgment and monitor AI quality and behavioral drift in production.
- Identify failure sources across models, prompts, retrieval, data quality, tools, agent logic, orchestration, and infrastructure.
- Establish evaluation-driven development practices and mentor other engineers.
Requirements
- 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
- Strong Python engineering skills and experience building production systems.
- Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
- Experience with multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms.
- Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating.
- Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering.
- Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
- Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders.
- Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment.
- Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis.
- Strong analytical, systems-thinking, and communication skills.
Benefits
- Comprehensive health, dental, and vision coverage.
- Generous parental leave policies, an Employee Assistance Program, wellness programs, a 401K with company matching, a Lifestyle Spending Account, and unlimited vacation.
- Hybrid work arrangement in or near San Francisco, California.
- The role is eligible for a variable plan and other comprehensive benefits in addition to base salary.
Tech Stack
MLflowPython