Truist Financial Corporation

Site Reliability Engineer

Truist Financial Corporation
Apply
2 hours ago
Tokyo, JapanSenior

Responsibilities

  • Lead major and high-severity incident response, root-cause diagnosis, cross-team resolution, and problem management to closure.
  • Architect and deliver automation that reduces toil and MTTR while improving service resilience.
  • Implement intelligent alerting, anomaly detection, event correlation, SLO/SLI adoption, and operational metrics.
  • Standardize observability across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
  • Ensure operational readiness through resiliency testing, chaos engineering, and failure-mode validation.
  • Develop and enforce runbooks, incident playbooks, escalation paths, recovery patterns, SRE frameworks, and maturity models.
  • Lead workshops, maturity assessments, enablement sessions, and Communities of Practice supporting enterprise reliability transformation.
  • Coach and mentor Associate, Professional, and Senior SREs while providing technical guidance, design reviews, and thought leadership.
  • Collaborate with delivery, architecture, security, risk, product, and technology teams to embed reliability into design and execution.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related field.
  • At least 7 years of professional software development experience.
  • At least 7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations is preferred.
  • Deep knowledge of multiple programming languages, software architecture, design principles, software development lifecycle, testing, deployment, and security practices.
  • Expertise with distributed systems, Kubernetes, cloud-native architectures, container orchestration, and operational tooling.
  • Proficiency with Python, Go, PowerShell, and Ansible.
  • Strong understanding of observability platforms including Splunk and Dynatrace, event-driven monitoring, networking, and Linux/Unix internals.
  • Proven leadership in major incident management, cross-team technical coordination, and executive-level communication during critical incidents.
  • Demonstrated ability to influence technical roadmaps and drive reliability best practices.
  • Preferred qualifications include an advanced technical degree, CSDP or equivalent certification, financial-services or regulated-industry experience, SRE transformation experience, chaos engineering, resilience assessments, service failure modeling, hybrid- and multi-cloud operations, and Center for Enablement or Communities of Practice experience.

Benefits

  • Regular teammates working 20 or more hours per week are eligible for medical, dental, vision, life, disability, accidental death and dismemberment, tax-preferred savings, and 401(k) benefits, subject to plan eligibility.
  • Eligible employees receive at least 10 days of vacation, 10 sick days, and paid holidays in the first year, prorated where applicable.
  • Depending on position and division, benefits may include a defined benefit pension plan, restricted stock units, and deferred compensation.
  • This is a regular, non-temporary position requiring onsite work Monday through Friday in Charlotte, North Carolina; Raleigh, North Carolina; or Atlanta, Georgia.

Tech Stack

AnsibleGoKubernetesLinuxPowerShellPythonSplunk

Categories

DevOpsSite Reliability
Truist Financial Corporation

About Truist Financial Corporation

10,000+ employees
Contact me