Jack Henry & Associates

Senior Software Engineer: Site Reliability Engineering

Jack Henry & Associates
Apply
27 days ago
Remote, Worldwide +3 moreSenior

Responsibilities

  • Drive reliability and performance for public cloud and internal server infrastructure environments.
  • Design and implement SRE practices, including SLOs, SLIs, error budgets, proactive system health monitoring, and toil reduction.
  • Develop Infrastructure as Code, configuration management, automation, and self-healing tooling using Terraform, Ansible, and GitHub.
  • Lead the redesign, migration, and redeployment of on-premises services into Google Cloud.
  • Establish comprehensive observability, monitoring, and alerting systems across hybrid environments.
  • Lead blameless post-mortems and root cause analyses for critical incidents.
  • Automate patching, vulnerability remediation, and compliance activities across infrastructure environments.
  • Partner with DevOps and development teams to embed reliability practices throughout the software development lifecycle.
  • Maintain SRE processes, operational runbooks, configuration documentation, and other technical documentation.
  • Participate in an on-call rotation approximately every 7–8 weeks.

Requirements

  • At least 6 years of experience in cloud and hybrid datacenter operations focused on Infrastructure as Code and Site Reliability Engineering.
  • Proficiency with GCP, AWS, and/or Azure.
  • Proficiency with GitOps, Terraform, and Ansible in a CI/CD pipeline.
  • Experience with PowerShell, Python, or GoLang.
  • Strong understanding of Linux/POSIX and Windows system administration, networking, and firewalls.
  • Understanding of security best practices and compliance standards including CIS, NIST, and PCI.
  • Ability to participate in an on-call rotation every 7–8 weeks.
  • A bachelor’s degree in Computer Science, Information Technology, or Engineering is preferred.
  • Relevant industry certifications, including Google Associate Cloud Engineer or Google Cloud Architect, are preferred.
  • Proficiency with ArgoCD and GitOps is preferred.
  • Familiarity with SQL and NoSQL databases is preferred.
  • Experience with OpenTelemetry tooling and alerting systems such as Prometheus, Grafana, and ELK Stack is preferred.
  • Knowledge of SRE principles including SLOs, SLIs, automation, toil reduction, and root cause analysis is preferred.

Benefits

  • Remote work is available within the United States, except California.
  • The role may require an onsite interview or in-person onboarding to verify identity.
  • The company offers comprehensive benefits supporting associates’ physical, mental, and financial health.
  • The position is ineligible for immigration sponsorship or support, including H-1B and PERM sponsorship.

Tech Stack

AnsibleAWSAzureGoGoogle Cloud PlatformGrafanaLinuxPowerShellPrometheusPythonSQLTerraformWindows

Categories

DevOpsSite Reliability
Contact me