
Senior Site Reliability Engineer
Crunchyroll, LLC3 hours ago
Hyderābād, IndiaStaff+
Responsibilities
- Define and improve platform reliability, availability, and performance using SLIs, SLOs, and error budgets.
- Establish incident management, root cause analysis, postmortem, and service ownership practices.
- Build monitoring, logging, tracing, alerting, automation, self-service, and self-healing capabilities.
- Design and optimize cloud-native infrastructure, platform standards, deployment automation, and infrastructure as code.
- Perform capacity planning, performance optimization, disaster recovery, backup, and business continuity work.
- Partner with security teams on SecOps, vulnerability remediation, penetration testing findings, cloud security, container security, and Kubernetes security.
- Collaborate with Engineering, Data, Product, Infrastructure, and Security teams on reliability, scalability, and security initiatives.
Requirements
- 8+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or related disciplines operating and scaling production-critical systems.
- Strong experience deploying, operating, and troubleshooting Kubernetes and GCP environments at scale.
- Excellent infrastructure as code experience, preferably with Terraform.
- Understanding of Linux systems administration, networking fundamentals, and distributed systems.
- Proficiency in one or more of Go, Python, Java, or Shell.
- Experience with observability and monitoring platforms such as Prometheus, Grafana, OpenTelemetry, or Datadog.
- Working knowledge of incident management, service reliability, capacity planning, performance optimization, SLIs, and SLOs.
- Experience with cloud and platform security, container security, Kubernetes security, vulnerability remediation, and secure infrastructure operations.
- Familiarity with OWASP Top 10, Identity and Access Management, secrets management, SSDLC, and security-by-design principles.
- Strong communication, collaboration, and problem-solving skills.
Benefits
- Best-in-class medical, dental, and vision private insurance coverage.
- 24/7 counseling and mental health sessions through the Employee Assistance Program.
- Free premium Crunchyroll access and professional development opportunities.
- Paid parental leave of up to 26 weeks for birthing parents and up to 12 weeks for non-birthing parents.
- Hybrid work schedule, paid time off, flex time off, five Yasumi Days, summer half-day Fridays, and winter break.
Tech Stack
Categories
DevOpsSite Reliability