
Site Reliability Engineering
Truist Financial Corporation2 hours ago
Tokyo, JapanStaff+
Responsibilities
- Lead major and high-severity incident responses, diagnose root causes, and coordinate multi-team technical resolution.
- Drive problem management to closure and implement systemic fixes for recurring operational risks.
- Architect and deliver automation that reduces toil and MTTR while improving service resilience.
- Implement intelligent alerting, anomaly detection, event correlation, and AIOps-enabled operational improvements.
- Establish SLO/SLI adoption, observability standards, dashboards, telemetry coverage, and operational KPIs.
- Lead resiliency testing, chaos engineering, failure-mode validation, maturity assessments, workshops, and enablement sessions.
- Develop and enforce incident playbooks, escalation paths, runbooks, response procedures, automated recovery patterns, and enterprise SRE frameworks.
- Provide technical leadership, coaching, mentoring, design guidance, and knowledge sharing for SRE and engineering teams.
- Coordinate with Delivery, Architecture, Security, Risk, product, and technology teams to influence roadmaps and reliability practices.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related field.
- Minimum of 7 years of professional experience in software development.
- At least 7 years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Infrastructure Operations is preferred.
- Deep knowledge of multiple programming languages, software architecture, design principles, software development lifecycle, testing, deployment, and security practices.
- Hands-on experience with distributed systems, Kubernetes, cloud-native operational tooling, automation scripting, observability platforms, networking, Linux/Unix internals, and modern infrastructure patterns.
- Proven leadership in major incident management, cross-team technical coordination, executive-level incident communication, and technical roadmap influence.
- Preferred qualifications include an advanced technical degree, CSDP or equivalent certification, financial services or regulated-industry experience, SRE transformation experience, chaos engineering, resilience assessments, failure-mode modeling, hybrid- or multi-cloud operations, and Center for Enablement or Communities of Practice experience.
- Strong communication, mentoring, coaching, technical judgment, and change leadership skills.
Benefits
- Regular employees working 20 or more hours per week may be eligible for medical, dental, vision, life, disability, accidental death and dismemberment, tax-preferred savings, and 401(k) benefits.
- Eligible employees receive at least 10 days of vacation, 10 sick days, and paid holidays during the first year, prorated as applicable.
- Depending on position and division, benefits may include a defined benefit pension plan, restricted stock units, and deferred compensation.
- Regular onsite work is required Monday through Friday in Charlotte, North Carolina; Raleigh, North Carolina; or Atlanta, Georgia.
Tech Stack
Categories
DevOpsSite Reliability