
Staff Site Reliability Engineer
Crunchyroll, LLC10 hours ago
Base Salary
$233k - $292k/yr
Responsibilities
- Define and improve platform reliability, availability, performance, SLIs, SLOs, and error budgets.
- Lead incident management, root cause analysis, postmortems, service ownership, and operational excellence practices.
- Build monitoring, logging, tracing, alerting, automation, self-service, and self-healing capabilities.
- Design and optimize cloud-native infrastructure, infrastructure as code, deployment automation, platform scalability, and capacity planning.
- Develop and validate disaster recovery, backup, resilience, and business continuity strategies.
- Own vulnerability triage and remediation across infrastructure, platforms, containers, and applications.
- Support penetration-testing scoping and remediate resulting security findings.
- Implement secure cloud, container, and Kubernetes environments using least privilege, defense in depth, and Zero Trust principles.
- Collaborate cross-functionally, influence architecture, drive reliability and security initiatives, and mentor engineers.
Requirements
- 12+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or related disciplines operating and scaling production-critical systems.
- Deep expertise in Kubernetes and GCP, including designing, deploying, and operating highly available cloud-native platforms.
- Strong infrastructure as code experience, preferably with Terraform.
- Experience with Linux systems administration, networking, distributed systems, and complex production troubleshooting.
- Proficiency in one or more of Go, Python, Java, or Shell for automation and platform engineering.
- Hands-on experience with Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent observability solutions.
- Expertise in incident management, service reliability, capacity planning, performance optimization, SLIs, SLOs, and error budgets.
- Understanding of cloud and platform security, container security, Kubernetes security, vulnerability management, and secure infrastructure operations.
- Knowledge of OWASP Top 10, Identity and Access Management, secrets management, Secure Software Development Lifecycle, and security-by-design principles is preferred.
- Strong collaboration, communication, technical leadership, architectural influence, cross-functional initiative leadership, and mentoring skills.
Benefits
- Compensation package including salary and annual performance bonus potential.
- Flexible time off policies.
- Medical, dental, vision, short-term disability, long-term disability, and life insurance.
- Health Savings Account and health care and dependent care Flexible Spending Account programs.
- 401(k) plan with employer match.
- Employer-paid commuter benefit.
- New-parent support program.
- Pet insurance and some pet-friendly offices.
- The posting is associated with a hybrid work arrangement.
Tech Stack
Categories
DevOpsSite Reliability
About Crunchyroll, LLC
Crunchyroll, LLC runs a global anime-focused streaming service with subscription and ad-supported tiers, plus manga, games, theatrical releases, and an e-commerce store for merchandise. It serves fans in 200+ countries and territories and monetizes through subscriptions, advertising, and retail. Founded in 2006 and headquartered in Los Angeles, it is owned by Sony through a joint venture between Sony Pictures Entertainment and Aniplex.