
Senior Systems Reliability Engineer
Electric Reliability Council of Texas (ERCOT)1 month ago
Taylor, TX, USASenior
Base Salary
$109k - $150k/yr
Responsibilities
- Build production reliability software, automated remediation systems, self-healing infrastructure components, and operational automation frameworks.
- Own SLO and error budget governance, SLI instrumentation, alert quality, observability architecture, and production readiness standards.
- Lead high-severity incident response, dual-datacenter failover execution, root-cause analysis, post-mortems, and remediation tracking.
- Design and implement chaos engineering programs, resilience tests, CI/CD reliability gates, canary analysis, progressive delivery validation, and rollback triggers.
- Diagnose Java/Spring Boot and JVM performance issues involving heap pressure, garbage collection, thread pools, connection leaks, and class loading.
- Design and operate Kubernetes and OpenShift workloads and build infrastructure-as-code using Terraform, Ansible/AAP, and Azure Resource Manager templates.
- Implement NERC/CIP compliance controls, configuration drift detection, access validation, audit evidence, and regulatory audit preparation.
- Mentor Systems Reliability Specialist I and II employees, lead design reviews, establish reliability engineering standards, and contribute shared tooling and technical documentation.
- Participate in a 24/7 on-call rotation and coordinate team delivery and on-call activities as needed.
Requirements
- Minimum 5 years of progressive experience in systems reliability, software with an SRE focus, or a closely related discipline.
- Bachelor's degree in Computer Science, Software Engineering, MIS, or a related field; equivalent education and experience providing comparable knowledge is accepted.
- Production experience building and operating SLO frameworks, observability platforms, chaos engineering programs, and automated remediation tooling.
- Proficiency in Python and Java, with experience writing production-quality reliability tooling and automation; Bash and Linux experience required.
- Deep Java, Spring Boot, and JVM performance expertise, including heap analysis, garbage collection tuning, thread profiling, and application instrumentation.
- Experience with Kubernetes, OpenShift, CI pipelines, Open telemetry, infrastructure-as-code, and high-severity incident response.
- Expertise with Grafana LGTM tooling, including Loki, Grafana, Tempo, and Mimir, plus Dynatrace and Splunk.
- Experience with Terraform, Ansible/AAP, or equivalent; networking knowledge and experience with container networking troubleshooting.
- Experience leading blameless post-mortem programs; dual-datacenter or hybrid-cloud reliability architecture experience preferred.
- NERC/CIP compliance experience involving control implementation, audit preparation, and regulatory engagement preferred.
- Master's degree in Computer Science, Software Engineering, or a related field is preferred.
- Azure certification, Certified Kubernetes Administrator certification, and ITIL Foundation or Managing Professional certification are preferred.
Benefits
- Hybrid work location in Taylor, Texas, with two days per week onsite.
- Expected salary range of $109,000 to $150,000 per year.
- 24/7 on-call rotation participation is part of the role.
Tech Stack
AnsibleApache KafkaAzureBashDatadogGrafanaJavaKubernetesLinuxOpenShiftPostgreSQLPythonSOAPSplunkSpring BootTerraform
Categories
Site Reliability