
Lead Site Reliability Engineer
RB (Reckitt Benckiser)10 days ago
Richmond, VA, USA or San Francisco, CA, USAStaff+
Base Salary
$147k - $234k/yr
Responsibilities
- Design, implement, and maintain highly available, scalable, and resilient cloud systems.
- Establish and monitor SLIs, SLOs, and SLAs and lead incident response, root cause analysis, and preventive improvements.
- Develop disaster recovery and business continuity plans for critical systems.
- Architect and manage AWS infrastructure using Terraform and optimize cloud resource utilization and cost.
- Automate deployment pipelines, monitoring, and operational workflows.
- Build internal tools and services, develop monitoring and alerting solutions, and improve observability.
- Collaborate with software engineering teams on reliability practices, system design, code reviews, and architectural decisions.
- Integrate security practices into pipelines, maintain infrastructure and application controls, and conduct security assessments and vulnerability management.
- Mentor junior SREs, promote SRE culture, document runbooks and technical specifications, and drive technical initiatives.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- At least 7 years of experience in Site Reliability Engineering, DevOps, or related roles.
- At least 3 years in a lead or senior technical position.
- Strong proficiency in Java, Python, and Node.js, with experience writing clean, maintainable, and testable code.
- Experience with large-scale production systems, microservices, distributed systems, serverless architectures, and event-driven systems.
- Extensive AWS experience and advanced Terraform skills.
- Expert-level GitLab knowledge, including CI/CD pipelines, runners, and GitOps.
- Experience with Docker, Kubernetes or ECS, configuration management, monitoring, log aggregation, and distributed tracing.
- Hands-on SAST experience and knowledge of DAST, OWASP Top 10, compliance frameworks, secrets management, and IAM.
- Experience with on-call rotations and incident management.
- Experience with chaos engineering, multi-cloud or hybrid cloud environments, and Agile/Scrum methodologies.
- AWS Solutions Architect or DevOps Engineer certification is preferred.
- Working knowledge of LLMs and agentic applications is a plus.
Benefits
- Full-time, regular, exempt position.
- The selected candidate must reside within a reasonable commuting distance and work full-time onsite.
- Eligible hiring locations are Richmond, Virginia, and San Francisco, California; San Francisco and Richmond are preferred locations for the System IT team.
- Reasonable accommodations are available for applicants and employees with disabilities.
- The Federal Reserve Bank is an equal opportunity employer.
Categories
DevOpsSite Reliability