
Staff Site Reliability Engineer
SentinelOne6 hours ago
Prague, CzechiaStaff+
Responsibilities
- Drive or participate in production incident management, rapid recovery, troubleshooting, and root-cause analysis.
- Lead incident postmortems and document findings and follow-up actions to prevent recurrence.
- Improve and optimize observability strategies, monitoring solutions, and alerting accuracy.
- Define and refine SLOs, SLIs, and SLAs aligned with business and customer objectives.
- Automate repetitive operational tasks and improve incident-response workflows.
- Support reliability initiatives for a multi-cloud platform processing large-scale data and security events.
- Mentor peers and junior engineers in incident response, troubleshooting, and reliability practices.
Requirements
- At least 5 years of experience in Site Reliability Engineering, DevOps, or a related field in cloud-native environments.
- Experience troubleshooting complex issues under pressure and improving established runbooks.
- Experience with Kubernetes and container orchestration.
- Experience with observability stacks such as Prometheus, Grafana, ELK, and OpenTelemetry.
- Proficiency in Python or Go and Bash scripting.
- Familiarity with CI/CD pipelines, including ArgoCD and GitHub Actions, and DevOps practices.
- Excellent communication skills and demonstrated ability to mentor peers in reliability practices.
Benefits
- Restricted Stock Units, an Employee Stock Purchase Plan, and a performance-based bonus.
- Flexible time off, standard five weeks of PTO, wellness days, paid short-term sick/nursing leave, and gender-neutral parental leave.
- Private medical care, life and disability insurance, pension contribution, travel medical insurance, and an Employee Assistance Program.
- Global home office, meal, wellbeing, and MultiSport allowances, plus access to the Wellness Coach app.
- Hybrid work in Prague or Brno, or remote work in Czechia/Slovakia; Prague-based employees work from the office at least two days per week.
Tech Stack
Categories
Site Reliability
About SentinelOne
SentinelOne builds an AI-driven cybersecurity platform that provides endpoint protection, detection and response, XDR, cloud, and identity security for enterprises. Its Singularity platform is sold as subscription software and managed services to secure laptops, servers, containers, and cloud workloads with automated detection and remediation. Founded in 2013 and headquartered in Mountain View, California, SentinelOne is a public company listed on the NYSE following its 2021 IPO.