about 3 hours ago
Responsibilities
- Build a deployment health dashboard with live model health metrics, monitoring, and proactive incident alerting.
- Create session-to-model lineage connecting affected sessions and user profiles to model IDs, impact scope, and root causes.
- Build a feedback improvement flywheel using Intercom tickets, CSAT feedback, and AI agents to identify related sessions.
- Retrieve execution traces for flagged sessions and generate summaries usable by non-engineering stakeholders.
- Filter high-value feedback into training data, ship improved models, and monitor post-deployment performance.
- Measure revenue against inference cost for each model deployment and support model selection with unit economics.
- Establish mature LLMOps practices for tracing, evaluation, and model incident response.
- Partner with researchers and engineers working on ASR, note generation, Evidence, and Dictate models to embed observability.
Requirements
- 2–3 years of hands-on LLMOps experience building observability, tracing, evaluation, and feedback systems for production LLMs.
- Experience at an AI company operating at or beyond Heidi’s maturity, most likely in the US or China.
- Proven ability to build monitoring and alerting systems using Datadog or similar tools.
- Experience implementing distributed tracing across multi-step LLM pipelines and scalable session and event data models.
- Hands-on experience building with LLMs, including agents that triage feedback and match tickets to sessions.
- Experience joining cost and revenue data into actionable per-model unit economics.
- Ability to take ambiguous mandates, design systems, and operate them in production.
- No PhD is required; demonstrated shipped work is valued.
- Preferred background in backend engineering, data platforms, or ML infrastructure.
- Preferred experience integrating Intercom with engineering systems.
- Healthcare or other regulated, safety-critical domain experience is preferred.
Benefits
- Annual learning and development budget
- Monthly health and wellness allowance
- Home office budget
- 26 weeks of paid primary parental leave
- 18 weeks of paid secondary parental leave
- Fertility support
- Four weeks of work from anywhere per year
- Equity
- Flexible self-managed schedule focused on outcomes
- Sustainable performance and mental health support
