3 months ago
Pleasant Grove, UT, USAStaff+
Responsibilities
- Architect and implement enterprise-scale infrastructure for web, mobile, backend, and data engineering teams.
- Define reliability standards, architectural patterns, and engineering best practices across the organization.
- Lead performance optimization, monitoring, observability, and alerting initiatives.
- Build automation for infrastructure provisioning, configuration management, and deployments.
- Establish incident response frameworks, playbooks, escalation procedures, and post-incident improvement processes.
- Design fault-tolerant architectures, automated testing frameworks, and disaster recovery strategies.
- Provide technical leadership and mentorship while contributing to the strategic direction of the SRE practice.
Requirements
- 10+ years of extensive experience as a Site Reliability Engineer or similar role, with a record of architecting large-scale distributed systems.
- Expert-level proficiency in multiple programming languages, including Python, Go, or Node.js.
- Advanced experience architecting multi-region, highly available systems using AWS and GCP.
- Deep expertise administering and architecting Kubernetes, including large-scale clusters, custom controllers, and cluster performance optimization.
- Extensive experience with observability platforms, custom monitoring solutions, and sophisticated alerting strategies.
- Experience designing complex enterprise-scale IAM architectures.
- Distinguished expertise in Infrastructure as Code with Terraform, including custom provider development and multi-cloud deployments.
- Demonstrated ability to resolve critical production issues in complex, high-stakes environments.
Categories
Site Reliability
