CaaS Private Site Reliability Lead Engineer - Vice President
D.B. Group20 days ago
Cary, NC, USAStaff+
Base Salary
$125k - $185k/yr
Responsibilities
- Lead the reliability strategy for the US CaaS Private platform, including SLO frameworks, operational standards, and incident management maturity.
- Drive resilience improvements across observability, capacity planning, upgrade safety, disaster readiness, and operational automation.
- Lead troubleshooting of complex production management issues, identify root causes, and implement preventive fixes.
- Define service indicators, alert thresholds, dashboard standards, production readiness criteria, escalation paths, and postmortem follow-through.
- Develop automation and self-healing workflows to reduce manual intervention, improve recovery times, and strengthen platform supportability.
- Partner with engineering, operations, and application teams to ensure platform changes are measurable, supportable, and aligned with reliability objectives.
- Mentor junior and middle engineers and foster blameless learning, measurable reliability, and operational excellence.
- Influence platform architecture and roadmap decisions using incident, capacity, operational trend, and reliability data.
- Communicate strategic insights and practical recommendations to technical, non-technical, engineering, operations, and business stakeholders.
Requirements
- Extensive hands-on experience with bare-metal Kubernetes, Linux, distributed systems reliability, and production platform operations.
- Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root-cause analysis.
- Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements.
- Experience defining SLOs, service indicators, alert-quality standards, production-readiness practices, and escalation models.
- Ability to independently lead complex reliability improvements while partnering across engineering, operations, and application teams.
- Ability to use AI tools to improve productivity and optimize workflows while applying responsible and ethical judgment regarding data and AI outputs.
- Strong communication skills for explaining technical findings to technical and non-technical stakeholders.
- Sound operational judgment when balancing urgency, risk, and long-term platform stability during incidents.
- Experience mentoring engineers and improving operational culture across platform or infrastructure teams.
- Background in automation tools, infrastructure-as-code practices, capacity planning, disaster readiness, or cloud-native platform operations.
Benefits
- Hybrid working model with in-office and work-from-home flexibility; employees hired into this role are expected to work in the Cary, NC office in accordance with the Bank’s hybrid model.
- Generous vacation, personal, and volunteer days.
- Health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits.
- Educational resources, matching gift programs, volunteer programs, and Employee Resource Groups.
- Diverse and inclusive work environment with professional development opportunities.
Tech Stack
Categories
Site Reliability