
Site Reliability Engineer
Truist Financial Corporation2 hours ago
Tokyo, JapanSenior
Responsibilities
- Lead major and high-severity incident response, root-cause diagnosis, cross-team resolution, and problem management to closure.
- Architect and deliver automation that reduces toil and MTTR while improving service resilience.
- Implement intelligent alerting, anomaly detection, event correlation, SLO/SLI adoption, and operational metrics.
- Standardize observability across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
- Ensure operational readiness through resiliency testing, chaos engineering, and failure-mode validation.
- Develop and enforce runbooks, incident playbooks, escalation paths, recovery patterns, SRE frameworks, and maturity models.
- Lead workshops, maturity assessments, enablement sessions, and Communities of Practice supporting enterprise reliability transformation.
- Coach and mentor Associate, Professional, and Senior SREs while providing technical guidance, design reviews, and thought leadership.
- Collaborate with delivery, architecture, security, risk, product, and technology teams to embed reliability into design and execution.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related field.
- At least 7 years of professional software development experience.
- At least 7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations is preferred.
- Deep knowledge of multiple programming languages, software architecture, design principles, software development lifecycle, testing, deployment, and security practices.
- Expertise with distributed systems, Kubernetes, cloud-native architectures, container orchestration, and operational tooling.
- Proficiency with Python, Go, PowerShell, and Ansible.
- Strong understanding of observability platforms including Splunk and Dynatrace, event-driven monitoring, networking, and Linux/Unix internals.
- Proven leadership in major incident management, cross-team technical coordination, and executive-level communication during critical incidents.
- Demonstrated ability to influence technical roadmaps and drive reliability best practices.
- Preferred qualifications include an advanced technical degree, CSDP or equivalent certification, financial-services or regulated-industry experience, SRE transformation experience, chaos engineering, resilience assessments, service failure modeling, hybrid- and multi-cloud operations, and Center for Enablement or Communities of Practice experience.
Benefits
- Regular teammates working 20 or more hours per week are eligible for medical, dental, vision, life, disability, accidental death and dismemberment, tax-preferred savings, and 401(k) benefits, subject to plan eligibility.
- Eligible employees receive at least 10 days of vacation, 10 sick days, and paid holidays in the first year, prorated where applicable.
- Depending on position and division, benefits may include a defined benefit pension plan, restricted stock units, and deferred compensation.
- This is a regular, non-temporary position requiring onsite work Monday through Friday in Charlotte, North Carolina; Raleigh, North Carolina; or Atlanta, Georgia.
Tech Stack
Categories
DevOpsSite Reliability