
Site Reliability Engineer
Truist Financial Corporation2 hours ago
Tokyo, JapanSenior
Responsibilities
- Lead major and high-severity incident response, diagnose root causes, coordinate cross-team resolution, and drive problem management to closure.
- Architect and deliver automation that reduces toil and mean time to recovery while improving service resilience.
- Standardize observability practices across logs, metrics, traces, and events using platforms such as Dynatrace and Splunk.
- Guide adoption of SLOs and SLIs, intelligent alerting, anomaly detection, event correlation, resiliency testing, and chaos engineering.
- Design and implement scalable, secure, highly available software solutions and reliability improvements.
- Develop and enforce incident playbooks, escalation paths, runbooks, recovery patterns, frameworks, templates, and maturity models.
- Lead complex initiatives, technical design reviews, workshops, maturity assessments, and Communities of Practice across teams.
- Coach and mentor Associate, Professional, and Senior SREs while providing technical guidance and reviewing delegated work.
- Evaluate emerging technologies and influence technical roadmaps, reliability standards, and enterprise operational maturity.
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or a related field.
- At least 7 years of professional software development experience.
- Deep knowledge of multiple programming languages, software architecture, design principles, software development lifecycle, testing, deployment, and security practices.
- Expertise in cloud-native architectures, microservices, container orchestration, DevOps, distributed systems, and cloud-native operational tooling.
- Proficiency with Python, Go, PowerShell, and Ansible for automation and scripting.
- Strong familiarity with Kubernetes, Splunk, Dynatrace, networking, Linux/Unix internals, and modern infrastructure patterns.
- Proven leadership in major incident management, cross-team technical coordination, and executive-level communication during critical incidents.
- Preferred qualifications include an advanced technical degree, CSDP or equivalent certification, financial services or regulated-industry experience, SRE transformation experience, and familiarity with chaos engineering and hybrid- or multi-cloud operations.
Benefits
- Regular employees working at least 20 hours per week may receive medical, dental, vision, life, disability, accidental death and dismemberment, tax-preferred savings, and 401(k) benefits.
- Eligible employees receive vacation, sick days, and paid holidays, with amounts prorated based on employment details.
- Depending on the position and division, benefits may include a defined benefit pension plan, restricted stock units, and deferred compensation.
- The position requires onsite work Monday through Friday in Charlotte, North Carolina; Raleigh, North Carolina; or Atlanta, Georgia.
Tech Stack
Categories
Site Reliability