Accelerant

Senior SRE

Accelerant
Apply
2 months ago
Remote, United StatesSenior

Responsibilities

  • Own the end-to-end reliability roadmap, including SLIs, SLOs, error budgets, observability pipelines, and investment priorities.
  • Harden the financial data platform’s availability, performance, recoverability, deployment paths, straight-through processing, and failover.
  • Instrument target systems using push and pull ingestion patterns and monitor both service health and business KPIs.
  • Build on-call, alerting, severity, escalation, incident review, and blameless postmortem processes.
  • Route alerts through Datadog and Incident.io with ServiceNow as the system of record.
  • Develop automation, self-healing, data lineage, retention, capacity planning, and actionable reliability dashboards.
  • Validate reliability at scale with 5,000+ transactions before go-live.
  • Design and ship AI agents for incident triage, log analysis, and root-cause investigation using Cursor.
  • Deploy and govern SRE agents on the organization’s internal AI fabric.

Requirements

  • Proven experience designing, operating, and scaling reliable production systems.
  • Deep hands-on experience with Datadog, Prometheus/Grafana, and OpenTelemetry, including push and pull ingestion patterns.
  • Strong experience defining SLIs, SLOs, error budgets, and business-level KPIs.
  • Experience operating Snowflake and Fabric data platforms, MuleSoft integration layers, and D365, including F&O and/or Power Apps.
  • Hands-on incident management experience with Incident.io and ServiceNow, including on-call and postmortem practices.
  • Hands-on experience building with LLMs and AI coding assistants, particularly Cursor; agent-building and deployment experience is a bonus.
  • Ability to define and defend reliability strategy, targets, and operational metrics to engineering leadership and business stakeholders.
  • Strong communication skills and the ability to explain technical failures and reliability risks to varied audiences.
  • Ability to operate autonomously and take action in ambiguous, fast-changing environments.
  • Experience in insurance, fintech, or regulated financial services is preferred.
  • Familiarity with insurance and finance concepts, or willingness to learn them deeply, is preferred.
  • Experience with Red Panda/Kafka streaming and event pipelines, data lineage, retention, and auditability is preferred.
  • Working knowledge of chaos engineering, performance and load testing, and capacity planning is preferred.
  • Experience deploying AI agents on an internal AI platform or fabric, including governance, evaluation harnesses, and prompt/version management, is preferred.

Tech Stack

Apache KafkaAWSDatadogGrafanaPrometheusSnowflake

Categories

Site Reliability
Accelerant

About Accelerant

501-1,000 employees
Contact me