
Staff Site Reliability Engineer
Crunchyroll, LLC10 hours ago
Base Salary
$211k - $263k/yr
Responsibilities
- Define and improve platform reliability, availability, and performance using SLIs, SLOs, and error budgets.
- Establish incident management, root cause analysis, postmortem, and service ownership practices.
- Build monitoring, logging, tracing, alerting, automation, self-service, and self-healing capabilities.
- Design and optimize scalable cloud-native infrastructure, infrastructure as code, platform standards, and deployment automation.
- Lead capacity planning, performance optimization, disaster recovery, backup, and business continuity initiatives.
- Integrate security controls into platform operations, including vulnerability management, penetration-test remediation, and cloud, container, and Kubernetes security.
- Collaborate with Engineering, Data, Product, Infrastructure, and Security teams on reliability, scalability, and security initiatives.
- Mentor engineers and champion operational excellence, reliability, ownership, continuous improvement, and security awareness.
Requirements
- 12+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or related disciplines, operating and scaling production-critical systems.
- Deep expertise in Kubernetes and GCP, including designing, deploying, and operating highly available cloud-native platforms at scale.
- Strong infrastructure-as-code experience, preferably with Terraform.
- Experience with Linux systems administration, networking, distributed systems, and troubleshooting complex production issues.
- Proficiency in one or more of Go, Python, Java, or Shell for automation and platform engineering.
- Experience with observability platforms and practices including Prometheus, Grafana, OpenTelemetry, or Datadog.
- Expertise in incident management, service reliability, capacity planning, performance optimization, operational excellence, SLIs, SLOs, and error budgets.
- Understanding of cloud and platform security, container security, Kubernetes security, vulnerability management, and secure infrastructure operations.
- Knowledge of OWASP Top 10, Identity and Access Management, secrets management, Secure Software Development Lifecycle, and security-by-design principles is preferred.
- Strong collaboration, communication, technical leadership, architectural influence, cross-functional initiative leadership, and mentoring skills.
Benefits
- Salary plus performance bonus earning potential paid annually.
- Flexible time off policies.
- Medical, dental, vision, short-term disability, long-term disability, and life insurance.
- Health Savings Account and health care and dependent care Flexible Spending Account programs.
- 401(k) plan with employer match.
- Employer-paid commuter benefit, new-parent support program, and pet insurance.
- Some offices are pet friendly; the posting is marked hybrid.
Tech Stack
Categories
DevOpsSite Reliability
About Crunchyroll, LLC
Crunchyroll, LLC runs a global anime-focused streaming service with subscription and ad-supported tiers, plus manga, games, theatrical releases, and an e-commerce store for merchandise. It serves fans in 200+ countries and territories and monetizes through subscriptions, advertising, and retail. Founded in 2006 and headquartered in Los Angeles, it is owned by Sony through a joint venture between Sony Pictures Entertainment and Aniplex.