Fingerprint

Senior Site Reliability Engineer

Fingerprint
Apply
1 day ago
Remote, WorldwideSenior

Base Salary

$152k - $205k/yr

Responsibilities

  • Own the reliability of core production systems, including instrumentation, targets, operations, and production behavior under real traffic.
  • Define and maintain SLIs, SLOs, error budgets, dashboards, and alerts for critical system paths.
  • Improve alert quality, anomaly detection, correctness detection, and customer-visible issue detection.
  • Lead incident response, restore service, and produce actionable postmortems and follow-up work.
  • Build secure, resilient, and cost-efficient infrastructure with controls for failure modes, degradation, load shedding, and blast-radius containment.
  • Perform load testing, profiling, saturation analysis, capacity planning, and performance work.
  • Improve change safety through progressive delivery, automated rollback, pre-production signals, and safer deployment practices.
  • Manage infrastructure through Terraform or equivalent configuration and infrastructure code.
  • Design and ship software and developer-facing tooling that reduces operational toil.
  • Run staged game days and chaos exercises to identify system gaps and safe operating limits.
  • Partner with product engineering teams on production readiness, capacity, failure modes, rollback plans, runbooks, and on-call handoff.
  • Participate in and improve the on-call rotation through better runbooks, escalation, and pager-fatigue reduction.
  • Apply security review practices to personal and peer engineering work.
  • Mentor engineers through code review, pairing, and design feedback.

Requirements

  • 6–10 years of experience in SRE, production engineering, infrastructure, or backend engineering in primarily cloud-based environments, preferably AWS, with meaningful production ownership.
  • Demonstrated experience designing, shipping, operating, and owning a significant system end to end.
  • Hands-on experience defining and operating SLIs, SLOs, and error budgets.
  • Experience leading or serving as a primary responder for high-severity, customer-facing incidents and improving organizational learning from them.
  • Deep knowledge of distributed-system failure modes in high-throughput, low-latency environments, including saturation, cascading failures, retry storms, capacity limits, degradation, and load shedding.
  • Strong cloud infrastructure fundamentals covering networking, load balancing, containerization, EKS/Kubernetes, and distributed systems.
  • Hands-on experience managing infrastructure through Terraform or an equivalent tool.
  • Production programming ability in Go, Python, or a comparable language.
  • Fluency with observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry, including instrumenting systems directly.
  • Production experience operating Redis/ElastiCache, including cluster and shard management, failover, memory eviction, and scaling strategies.
  • Knowledge of source control, code review, comprehensive test coverage, and safe deployment practices.
  • High ownership and autonomy, including experience working with ambiguous requirements.
  • Strong written and verbal English communication for technical documents, reviews, incident updates, and postmortems.
  • Experience using AI tools for incident investigation, telemetry analysis, runbook writing, and tooling, with informed opinions about their use.

Benefits

  • 100% remote work arrangement with the ability to join from almost any country, subject to country restrictions.
  • Fingerprint does not sponsor visas, and teammates must be authorized to work from their home location.
  • Inclusive work environment with emphasis on respecting and valuing diverse experiences and backgrounds.

Tech Stack

AWSDatadogGoGrafanaKubernetesPrometheusPythonRedisTerraform

Categories

Site Reliability
Fingerprint

About Fingerprint

201-500 employees

Fingerprint builds a device intelligence platform and APIs that help fraud, risk, and security teams identify users, bots, and risky activity in real time. The company offers commercial SaaS for device fingerprinting and fraud prevention, alongside the open-source FingerprintJS project used by developers. Founded in 2020 and headquartered in Chicago, it is privately held and counts customers such as Dropbox, Booking.com, and Plaid.

Contact me