17 hours ago
Base Salary
$252k - $308k/yr
Responsibilities
- Define and evolve reliability standards covering SLIs, SLOs, error budgets, production readiness, observability, incident response, and resilience.
- Lead high-severity incidents as Incident Commander and improve the full incident lifecycle, including detection, triage, investigation, communication, postmortems, and corrective-action tracking.
- Design and implement AI-assisted workflows for alert correlation, context retrieval, runbook automation, root-cause exploration, postmortem drafting, and remediation recommendations.
- Improve on-call quality by reducing alert noise, automating repetitive investigations, and building tools that integrate operational context from monitoring, incident, communication, and deployment systems.
- Guide graceful degradation, failure isolation, capacity planning, and operational safety across EarnIn’s AWS environment.
- Coach engineers, lead design and incident reviews, and build reusable reliability documentation, tooling, and practices across product engineering teams.
Requirements
- 7+ years of experience in SRE, software engineering, or infrastructure engineering with increasing scope and cross-organizational influence.
- Demonstrated success improving reliability and operational excellence at scale using KPIs such as MTTR, MTTD, alert quality, incident recurrence, SLO attainment, on-call health, or corrective-action completion.
- Experience applying AI or LLMs to engineering or operational workflows such as alert triage, runbook automation, incident investigation, postmortems, remediation recommendations, knowledge retrieval, or agentic operations tooling.
- Significant expertise with SLIs, SLOs, error budgets, incident command, blameless postmortems, and recurrence prevention in large-scale distributed systems.
- Strong software engineering ability in Python, Go, or similar languages, with experience building tools and automation.
- Deep observability experience with Datadog, CloudWatch, OpenTelemetry, or similar platforms.
- Strong infrastructure-as-code and cloud infrastructure experience, including Terraform, Kubernetes, AWS, and safe reversible deployments.
- Practical experience with AI-assisted development tools such as Cursor, Claude Code, Copilot, or ChatGPT.
- Experience in fintech, regulated environments, SOC 2, PCI, FinOps, or high-scale cost and performance tradeoffs is preferred.
Benefits
- Base salary range is $252,000-$308,000 plus equity and benefits.
- Hybrid position based in Mountain View requiring in-office work two days per week.
Tech Stack
Categories
Site Reliability
About EarnIn
EarnIn builds a mobile app for earned wage access and related financial tools for U.S. workers living paycheck to paycheck. It lets users access a portion of accrued wages before payday, plus features like credit monitoring, automated savings, low-balance alerts, and a debit card, with no mandatory fees or interest. Founded in 2012 and headquartered in Mountain View, California, the privately held company issues certain banking products via Evolve Bank & Trust.