
Senior Software Engineer: Site Reliability Engineering
Jack Henry & Associates27 days ago
Remote, Worldwide +3 moreSenior
Responsibilities
- Drive reliability and performance for public cloud and internal server infrastructure environments.
- Design and implement SRE practices, including SLOs, SLIs, error budgets, proactive system health monitoring, and toil reduction.
- Develop Infrastructure as Code, configuration management, automation, and self-healing tooling using Terraform, Ansible, and GitHub.
- Lead the redesign, migration, and redeployment of on-premises services into Google Cloud.
- Establish comprehensive observability, monitoring, and alerting systems across hybrid environments.
- Lead blameless post-mortems and root cause analyses for critical incidents.
- Automate patching, vulnerability remediation, and compliance activities across infrastructure environments.
- Partner with DevOps and development teams to embed reliability practices throughout the software development lifecycle.
- Maintain SRE processes, operational runbooks, configuration documentation, and other technical documentation.
- Participate in an on-call rotation approximately every 7–8 weeks.
Requirements
- At least 6 years of experience in cloud and hybrid datacenter operations focused on Infrastructure as Code and Site Reliability Engineering.
- Proficiency with GCP, AWS, and/or Azure.
- Proficiency with GitOps, Terraform, and Ansible in a CI/CD pipeline.
- Experience with PowerShell, Python, or GoLang.
- Strong understanding of Linux/POSIX and Windows system administration, networking, and firewalls.
- Understanding of security best practices and compliance standards including CIS, NIST, and PCI.
- Ability to participate in an on-call rotation every 7–8 weeks.
- A bachelor’s degree in Computer Science, Information Technology, or Engineering is preferred.
- Relevant industry certifications, including Google Associate Cloud Engineer or Google Cloud Architect, are preferred.
- Proficiency with ArgoCD and GitOps is preferred.
- Familiarity with SQL and NoSQL databases is preferred.
- Experience with OpenTelemetry tooling and alerting systems such as Prometheus, Grafana, and ELK Stack is preferred.
- Knowledge of SRE principles including SLOs, SLIs, automation, toil reduction, and root cause analysis is preferred.
Benefits
- Remote work is available within the United States, except California.
- The role may require an onsite interview or in-person onboarding to verify identity.
- The company offers comprehensive benefits supporting associates’ physical, mental, and financial health.
- The position is ineligible for immigration sponsorship or support, including H-1B and PERM sponsorship.