D.B. Group

CaaS Private Site Reliability Lead Engineer - Vice President

D.B. Group
Apply
20 days ago
Cary, NC, USAStaff+

Base Salary

$125k - $185k/yr

Responsibilities

  • Lead the reliability strategy for the US CaaS Private platform, including SLO frameworks, operational standards, and incident management maturity.
  • Drive resilience improvements across observability, capacity planning, upgrade safety, disaster readiness, and operational automation.
  • Lead troubleshooting of complex production management issues, identify root causes, and implement preventive fixes.
  • Define service indicators, alert thresholds, dashboard standards, production readiness criteria, escalation paths, and postmortem follow-through.
  • Develop automation and self-healing workflows to reduce manual intervention, improve recovery times, and strengthen platform supportability.
  • Partner with engineering, operations, and application teams to ensure platform changes are measurable, supportable, and aligned with reliability objectives.
  • Mentor junior and middle engineers and foster blameless learning, measurable reliability, and operational excellence.
  • Influence platform architecture and roadmap decisions using incident, capacity, operational trend, and reliability data.
  • Communicate strategic insights and practical recommendations to technical, non-technical, engineering, operations, and business stakeholders.

Requirements

  • Extensive hands-on experience with bare-metal Kubernetes, Linux, distributed systems reliability, and production platform operations.
  • Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root-cause analysis.
  • Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements.
  • Experience defining SLOs, service indicators, alert-quality standards, production-readiness practices, and escalation models.
  • Ability to independently lead complex reliability improvements while partnering across engineering, operations, and application teams.
  • Ability to use AI tools to improve productivity and optimize workflows while applying responsible and ethical judgment regarding data and AI outputs.
  • Strong communication skills for explaining technical findings to technical and non-technical stakeholders.
  • Sound operational judgment when balancing urgency, risk, and long-term platform stability during incidents.
  • Experience mentoring engineers and improving operational culture across platform or infrastructure teams.
  • Background in automation tools, infrastructure-as-code practices, capacity planning, disaster readiness, or cloud-native platform operations.

Benefits

  • Hybrid working model with in-office and work-from-home flexibility; employees hired into this role are expected to work in the Cary, NC office in accordance with the Bank’s hybrid model.
  • Generous vacation, personal, and volunteer days.
  • Health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits.
  • Educational resources, matching gift programs, volunteer programs, and Employee Resource Groups.
  • Diverse and inclusive work environment with professional development opportunities.

Tech Stack

Categories

Site Reliability
D.B. Group

About D.B. Group

501-1,000 employees
Contact me