
Senior SRE
Accelerant2 months ago
Remote, United StatesSenior
Responsibilities
- Own the end-to-end reliability roadmap, including SLIs, SLOs, error budgets, observability pipelines, and investment priorities.
- Harden the financial data platform’s availability, performance, recoverability, deployment paths, straight-through processing, and failover.
- Instrument target systems using push and pull ingestion patterns and monitor both service health and business KPIs.
- Build on-call, alerting, severity, escalation, incident review, and blameless postmortem processes.
- Route alerts through Datadog and Incident.io with ServiceNow as the system of record.
- Develop automation, self-healing, data lineage, retention, capacity planning, and actionable reliability dashboards.
- Validate reliability at scale with 5,000+ transactions before go-live.
- Design and ship AI agents for incident triage, log analysis, and root-cause investigation using Cursor.
- Deploy and govern SRE agents on the organization’s internal AI fabric.
Requirements
- Proven experience designing, operating, and scaling reliable production systems.
- Deep hands-on experience with Datadog, Prometheus/Grafana, and OpenTelemetry, including push and pull ingestion patterns.
- Strong experience defining SLIs, SLOs, error budgets, and business-level KPIs.
- Experience operating Snowflake and Fabric data platforms, MuleSoft integration layers, and D365, including F&O and/or Power Apps.
- Hands-on incident management experience with Incident.io and ServiceNow, including on-call and postmortem practices.
- Hands-on experience building with LLMs and AI coding assistants, particularly Cursor; agent-building and deployment experience is a bonus.
- Ability to define and defend reliability strategy, targets, and operational metrics to engineering leadership and business stakeholders.
- Strong communication skills and the ability to explain technical failures and reliability risks to varied audiences.
- Ability to operate autonomously and take action in ambiguous, fast-changing environments.
- Experience in insurance, fintech, or regulated financial services is preferred.
- Familiarity with insurance and finance concepts, or willingness to learn them deeply, is preferred.
- Experience with Red Panda/Kafka streaming and event pipelines, data lineage, retention, and auditability is preferred.
- Working knowledge of chaos engineering, performance and load testing, and capacity planning is preferred.
- Experience deploying AI agents on an internal AI platform or fabric, including governance, evaluation harnesses, and prompt/version management, is preferred.
Tech Stack
Categories
Site Reliability