LogicMonitor

Sr. Evaluation Engineer

LogicMonitor
Apply
17 days ago

Base Salary

$158k - $218k/yr

Responsibilities

  • Define quality metrics for incident diagnostics, root-cause analysis, alert correlation, grounding, tool use, safety, and operational usefulness.
  • Build Python-based offline and online evaluation pipelines integrated with experimentation, model selection, prompt iteration, CI/CD, and release gating.
  • Create and maintain golden datasets, regression suites, customer-specific scenarios, adversarial test cases, and evaluation-driven safeguards.
  • Design step-level and trajectory-level evaluations for multi-step and multi-agent workflows, including planning, retrieval, evidence use, tool selection, and escalation decisions.
  • Evaluate AI capabilities such as incident investigation, on-call assistance, operational question answering, change-impact analysis, remediation recommendations, and automated resolution.
  • Design and calibrate LLM-based graders against expert human judgment and monitor AI quality and behavioral drift in production.
  • Identify failure sources across models, prompts, retrieval, data quality, tools, agent logic, orchestration, and infrastructure.
  • Establish evaluation-driven development practices and mentor other engineers.

Requirements

  • 5+ years of experience in software engineering, machine learning, applied AI, or a related field.
  • Strong Python engineering skills and experience building production systems.
  • Hands-on experience with AI evaluation, experimentation, testing, and quality frameworks.
  • Experience with multiple LLM and agent evaluation frameworks, such as LangSmith, Arize Phoenix, Braintrust, DeepEval, Ragas, TruLens, OpenAI Evals, MLflow, or comparable platforms.
  • Ability to select, customize, and integrate evaluation frameworks for offline testing, online monitoring, regression analysis, experimentation, model and prompt comparison, and release gating.
  • Strong understanding of LLMs, agents, retrieval-augmented generation, prompt engineering, tool calling, and context engineering.
  • Experience evaluating non-deterministic, multi-step, or multi-agent AI systems.
  • Ability to translate human and domain-expert judgment into test cases, evaluation rubrics, scoring functions, and automated graders.
  • Experience with LLM-as-a-judge techniques, including grader design, calibration, reliability measurement, and alignment with expert human judgment.
  • Experience with regression testing, CI/CD, production monitoring, behavioral drift detection, and failure analysis.
  • Strong analytical, systems-thinking, and communication skills.

Benefits

  • Comprehensive health, dental, and vision coverage.
  • Generous parental leave policies, an Employee Assistance Program, wellness programs, a 401K with company matching, a Lifestyle Spending Account, and unlimited vacation.
  • Hybrid work arrangement in or near San Francisco, California.
  • The role is eligible for a variable plan and other comprehensive benefits in addition to base salary.

Tech Stack

MLflowPython
LogicMonitor

About LogicMonitor

1,001-5,000 employees
Contact me