Staff Site Reliability Engineer
BeyondTrustabout 9 hours ago
Responsibilities
- Design, scale, and maintain highly available, secure, and resilient systems across cloud and on-premises environments.
- Champion platform engineering initiatives to reduce cognitive load for software engineers.
- Own and optimize foundational services like API gateways and service meshes.
- Drive a culture of 'Everything as Code' for infrastructure and CI/CD pipelines.
- Standardize and secure modern CI/CD pipelines for rapid code deployments.
- Implement automated chaos engineering frameworks and disaster recovery simulations.
- Architect and mature the telemetry stack for system visibility.
- Define and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs).
- Partner with engineering leadership to define the long-term SRE strategy.
- Mentor engineers on reliability practices and build accessible documentation.
Requirements
- 7+ years of experience in SRE, DevOps, or Platform Engineering, with at least 2 years in a Senior or Staff level.
- Proven track record of managing automation and infrastructure across cloud and on-premises environments.
- Experience with Docker and Kubernetes, including cluster administration.
- Prior experience with release orchestration strategies like Canary or Blue Green models.
- Proficiency in at least one systems language such as Go, Java, or C#.
- Deep understanding of open telemetry and monitoring best practices.
- Familiarity with GitOps workflows and Infrastructure as Code best practices.