Staff Site Reliability Engineer
Fingerprint4 hours ago
Remote, WorldwideStaff+
Base Salary
$177k - $240k/yr
Responsibilities
- Define and operationalize SLIs, SLOs, error budgets, and reliability metrics for critical identification, events, server API, and client-agent request paths.
- Strengthen incident detection, response, communication, postmortems, follow-up, alert quality, anomaly detection, and escalation practices.
- Lead reliability reviews for high-risk changes and services, including production readiness, capacity, failure modes, and rollback planning.
- Introduce game days, chaos exercises, and deliberate failure testing in staging and production.
- Embed with teams to solve reliability problems, coach engineers, and establish durable on-call, runbook, change-safety, and production-readiness practices.
- Write production tooling, dashboards, and reference implementations and contribute hands-on during incidents and investigations.
- Lead AI adoption for incident investigation, postmortem analysis, runbook authoring, observability, and safe AI-assisted operations.
Requirements
- 10+ years of engineering experience, including 3+ years as an SRE, production engineer, or reliability-focused Staff engineer operating across multiple teams.
- Deep practical experience with SLI/SLO design, error budgets, incident leadership, high-severity customer-facing incidents, and postmortems.
- Hands-on experience with distributed-systems failure modes in high-throughput, low-latency environments, including saturation, cascading failures, retry storms, capacity limits, degradation, and load shedding.
- Fluency in Kubernetes, AWS, and modern observability tooling such as Datadog or equivalent.
- Ability to read and write production code in Go, TypeScript, or similar languages and work with infrastructure as code.
- Demonstrated ability to lead through influence, coach engineers, improve adoption, and communicate incidents, risks, and trade-offs to technical and executive audiences.
- Experience using AI tools for incident investigation, telemetry analysis, runbooks, postmortems, and tooling, plus judgment about safe AI-assisted operations.
- Preferred experience includes fraud detection, identity, payments, adversarial real-time systems, multi-region or cell-based architectures, Elasticsearch, Redis, DynamoDB, Kafka, and FinOps.
Benefits
- 100% remote work at a globally dispersed, all-remote company.
- Role is open across the United States and may be based anywhere company team members work, subject to country restrictions.
- The company does not sponsor visas; teammates must be authorized to work from their home location.
Tech Stack
Categories
Site Reliability
About Fingerprint
Fingerprint builds a device intelligence platform and APIs that help fraud, risk, and security teams identify users, bots, and risky activity in real time. The company offers commercial SaaS for device fingerprinting and fraud prevention, alongside the open-source FingerprintJS project used by developers. Founded in 2020 and headquartered in Chicago, it is privately held and counts customers such as Dropbox, Booking.com, and Plaid.