D.B. Group

CaaS Private Site Reliability Engineer - Assistant Vice President

D.B. Group
Apply
20 days ago
Cary, NC, USAStaff+

Base Salary

$100k - $153k/yr

Responsibilities

  • Define and continuously improve SLI/SLOs, alerting standards, and error budgets for the private CaaS platform and critical services.
  • Build and maintain observability across metrics, logs, alerts, and dashboards.
  • Lead or coordinate incident response, mitigation, communication, blameless postmortems, and follow-up actions.
  • Automate operational tasks and remediation workflows to reduce toil and improve recovery time.
  • Improve the reliability, upgrade safety, and operational readiness of Kubernetes clusters, ingress paths, service mesh components, node services, and platform dependencies.
  • Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, documentation, and best-practice adoption.

Requirements

  • Hands-on experience operating Kubernetes clusters on bare metal or private cloud environments and supporting platform services at scale.
  • Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or a related infrastructure role.
  • Strong Linux system administration and infrastructure scripting experience using Python, Ansible, and Bash.
  • Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, and dashboards.
  • Understanding of incident management, root cause analysis, operational readiness, networking, virtualization, containerization, and distributed systems behavior under failure.
  • Ability to use AI tools responsibly to improve productivity and solve business problems.
  • Preferred experience with Istio, Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails.
  • Familiarity with PostgreSQL, Kafka, MongoDB, or comparable stateful platform dependencies.
  • Knowledge of capacity planning, load testing, chaos testing, failure injection, alert tuning, and self-healing automation.
  • Experience with low-latency or regulated environments and constraints such as strict uptime, change control, compliance, time synchronization, deterministic performance, or SR-IOV.
  • Ability to read and understand Golang code when troubleshooting platform components.

Benefits

  • Hybrid working model with in-office and work-from-home flexibility; employees hired into this role are expected to work in the Cary, North Carolina office according to the bank’s hybrid model.
  • Generous vacation, personal, and volunteer days.
  • Health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits.
  • Employee Resource Groups, educational resources, matching gift programs, and volunteer programs.
  • Physical, emotional, and financial wellness benefits.

Tech Stack

AmbassadorAnsibleApache KafkaBashGoGrafanaIstioKubernetesLinuxMongoDBPostgreSQLPrometheusPythonSplunk

Categories

DevOpsSite Reliability
D.B. Group

About D.B. Group

501-1,000 employees
Contact me