CaaS Private Site Reliability Engineer - Assistant Vice President
D.B. Group20 days ago
Cary, NC, USAStaff+
Base Salary
$100k - $153k/yr
Responsibilities
- Define and continuously improve SLI/SLOs, alerting standards, and error budgets for the private CaaS platform and critical services.
- Build and maintain observability across metrics, logs, alerts, and dashboards.
- Lead or coordinate incident response, mitigation, communication, blameless postmortems, and follow-up actions.
- Automate operational tasks and remediation workflows to reduce toil and improve recovery time.
- Improve the reliability, upgrade safety, and operational readiness of Kubernetes clusters, ingress paths, service mesh components, node services, and platform dependencies.
- Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, documentation, and best-practice adoption.
Requirements
- Hands-on experience operating Kubernetes clusters on bare metal or private cloud environments and supporting platform services at scale.
- Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or a related infrastructure role.
- Strong Linux system administration and infrastructure scripting experience using Python, Ansible, and Bash.
- Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, and dashboards.
- Understanding of incident management, root cause analysis, operational readiness, networking, virtualization, containerization, and distributed systems behavior under failure.
- Ability to use AI tools responsibly to improve productivity and solve business problems.
- Preferred experience with Istio, Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails.
- Familiarity with PostgreSQL, Kafka, MongoDB, or comparable stateful platform dependencies.
- Knowledge of capacity planning, load testing, chaos testing, failure injection, alert tuning, and self-healing automation.
- Experience with low-latency or regulated environments and constraints such as strict uptime, change control, compliance, time synchronization, deterministic performance, or SR-IOV.
- Ability to read and understand Golang code when troubleshooting platform components.
Benefits
- Hybrid working model with in-office and work-from-home flexibility; employees hired into this role are expected to work in the Cary, North Carolina office according to the bank’s hybrid model.
- Generous vacation, personal, and volunteer days.
- Health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits.
- Employee Resource Groups, educational resources, matching gift programs, and volunteer programs.
- Physical, emotional, and financial wellness benefits.
Tech Stack
AmbassadorAnsibleApache KafkaBashGoGrafanaIstioKubernetesLinuxMongoDBPostgreSQLPrometheusPythonSplunk
Categories
DevOpsSite Reliability