
Senior Engineer, Insights & Reliability Engineering
Crunchyroll, LLC11 days ago
Mexico City, MexicoSenior
Responsibilities
- Build, operate, and evolve reliability tooling, telemetry workflows, alerting systems, dashboards, and automation.
- Transform high-volume device, payment, partner, product, and platform telemetry into operational signals for anomaly detection, debugging, alerting, and leadership visibility.
- Use SQL, logs, metrics, dashboards, and product context to identify anomalies, validate hypotheses, and uncover reliability patterns.
- Lead cross-team improvements to monitoring, alerting, signal quality, triage workflows, incident response, and operational automation.
- Participate in production incident response and investigate device-, payment-, and partner-related issues.
- Collaborate with service owners, engineering, product, data, analytics, and external technical partners to drive systemic fixes and postmortem learnings.
- Set technical direction for reliability tooling, telemetry quality, debugging workflows, and operational readiness.
- Communicate technical findings, reliability risks, customer impact, tradeoffs, and recommendations to technical and non-technical stakeholders.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field, and/or 8+ years of practical engineering experience.
- Strong coding skills, with Python preferred; experience with comparable modern programming languages is acceptable with the ability to ramp up in Python-based workflows.
- Strong SQL skills and experience with high-volume event, telemetry, product, or operational datasets.
- Experience building, testing, deploying, and maintaining production-quality software, internal tools, data workflows, automation, or observability systems.
- Experience with data-processing systems, analytics engineering workflows, reliability tooling, alerting systems, or operational insight platforms.
- Ability to debug ambiguous production problems using data, logs, metrics, dashboards, code, system behavior, and product context.
- Experience with platforms such as Databricks, Spark, Airflow, dbt, Datadog, Grafana, Tableau, Mixpanel, or Mux.
- Experience defining scalable operational processes, improving signal quality, reducing alert noise, and replacing manual workflows with durable automation.
- Experience in reliability engineering, SRE, incident response, support engineering, analytics engineering, data engineering, or operational tooling.
- Experience with consumer devices, payments, partner integrations, or distributed systems.
- Ability to influence stakeholders across Engineering, Product, Data, Analytics, and partner-facing teams.
- Experience leading ambiguous, cross-system initiatives and aligning teams around technical priorities, operational tradeoffs, and reliability improvements.
- Experience using AI-assisted engineering tools and agentic workflows with sound judgment around correctness, maintainability, security, and production safety.
- Clear communication skills and the ability to translate technical findings, reliability risks, and customer impact for broad audiences.
Benefits
- Hybrid work arrangement is indicated by the #LI-Hybrid designation.
Tech Stack
Categories
Site Reliability