Crunchyroll, LLC

Senior Engineer, Insights & Reliability Engineering

Crunchyroll, LLC
Apply
11 days ago
Mexico City, MexicoSenior

Responsibilities

  • Build, operate, and evolve reliability tooling, telemetry workflows, alerting systems, dashboards, and automation.
  • Transform high-volume device, payment, partner, product, and platform telemetry into operational signals for anomaly detection, debugging, alerting, and leadership visibility.
  • Use SQL, logs, metrics, dashboards, and product context to identify anomalies, validate hypotheses, and uncover reliability patterns.
  • Lead cross-team improvements to monitoring, alerting, signal quality, triage workflows, incident response, and operational automation.
  • Participate in production incident response and investigate device-, payment-, and partner-related issues.
  • Collaborate with service owners, engineering, product, data, analytics, and external technical partners to drive systemic fixes and postmortem learnings.
  • Set technical direction for reliability tooling, telemetry quality, debugging workflows, and operational readiness.
  • Communicate technical findings, reliability risks, customer impact, tradeoffs, and recommendations to technical and non-technical stakeholders.

Requirements

  • Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field, and/or 8+ years of practical engineering experience.
  • Strong coding skills, with Python preferred; experience with comparable modern programming languages is acceptable with the ability to ramp up in Python-based workflows.
  • Strong SQL skills and experience with high-volume event, telemetry, product, or operational datasets.
  • Experience building, testing, deploying, and maintaining production-quality software, internal tools, data workflows, automation, or observability systems.
  • Experience with data-processing systems, analytics engineering workflows, reliability tooling, alerting systems, or operational insight platforms.
  • Ability to debug ambiguous production problems using data, logs, metrics, dashboards, code, system behavior, and product context.
  • Experience with platforms such as Databricks, Spark, Airflow, dbt, Datadog, Grafana, Tableau, Mixpanel, or Mux.
  • Experience defining scalable operational processes, improving signal quality, reducing alert noise, and replacing manual workflows with durable automation.
  • Experience in reliability engineering, SRE, incident response, support engineering, analytics engineering, data engineering, or operational tooling.
  • Experience with consumer devices, payments, partner integrations, or distributed systems.
  • Ability to influence stakeholders across Engineering, Product, Data, Analytics, and partner-facing teams.
  • Experience leading ambiguous, cross-system initiatives and aligning teams around technical priorities, operational tradeoffs, and reliability improvements.
  • Experience using AI-assisted engineering tools and agentic workflows with sound judgment around correctness, maintainability, security, and production safety.
  • Clear communication skills and the ability to translate technical findings, reliability risks, and customer impact for broad audiences.

Benefits

  • Hybrid work arrangement is indicated by the #LI-Hybrid designation.

Tech Stack

Apache AirflowApache SparkDatabricksDatadogdbtGrafanaPythonSQL

Categories

Site Reliability
Crunchyroll, LLC

About Crunchyroll, LLC

1,001-5,000 employees
Contact me