Rocket Money

Senior Infrastructure Engineer, SRE

Rocket Money
Apply
2 hours ago
Remote, United States +3 moreSenior
H1B Sponsor

Base Salary

$150k - $185k/yr

Responsibilities

  • Build and improve the reliability and resiliency of production systems and services.
  • Establish and regularly review SLIs, SLOs, and error budgets for critical services and user journeys.
  • Own disaster recovery strategy, including recovery objectives, failover and restore paths, and regular exercises.
  • Partner with product engineering teams to improve service ownership, instrumentation, and user-focused metrics.
  • Evolve observability standards for metrics, tracing, logs, instrumentation, alert quality, and cost.
  • Strengthen incident response by tuning paging thresholds, maintaining runbooks, and tracking postmortem actions.
  • Contribute to cloud infrastructure build-outs, platform backlog work, automation, and a shared on-call rotation of one week every six weeks.
  • Lead or contribute to reliability and observability modernization, internal tooling, game days, chaos experiments, and disaster recovery exercises.

Requirements

  • 5+ years of hands-on cloud or infrastructure engineering experience, including substantial reliability and production operations work at scale.
  • Experience defining SLIs and SLOs for real production services and evaluating their operational impact.
  • Hands-on production experience with an observability platform; Datadog is strongly preferred.
  • Ability to write code in Python, Go, TypeScript, or a similar language for tooling, debugging, and automation.
  • Production Terraform experience and comfort working in AWS.
  • Experience building or operating disaster recovery plans, including recovery goals, failover and restore procedures, and drills.
  • Experience being on-call for services and improving alert quality.
  • Preference for enabling teams with paved roads and good defaults rather than mandates.
  • Preferred experience leading reliability or observability modernization projects and delivering their implementations.
  • Preferred experience building internal tooling, libraries, or instrumentation standards, running game days or chaos experiments, and reducing observability costs.

Benefits

  • Health, dental, and vision plans.
  • 401k matching.
  • Unlimited PTO.
  • Competitive pay, plus bonus and benefits.
  • Daily lunch, snacks, and coffee for in-office employees.
  • Commuter benefits for in-office employees.
  • Shared on-call rotation of one week every six weeks.

Categories

DevOpsSite Reliability
Rocket Money

About Rocket Money

201-500 employees

Rocket Money is a leading personal finance app with a mission to help everyone money. We offer over 5 million members a centralized destination to manage their finances, with valuable services that save both time and money — ultimately giving them a leg up on their financial journey. Members can manage their subscriptions, lower their bills, build budgets and automatically set aside money to reach their savings goals. At Rocket Money, we believe we can do our best work operating as a team. We are committed to using our experience at work to support our own and each other's personal growth. If you are someone who enjoys solving hard problems, thrives on collaboration and is passionate about our mission, let's talk. We hire from all industries and look for a diverse set of experiences in our talent. To see our open roles go to www.rocketmoney.com/careers.

Contact me