Metropolis

Staff Software Engineer, Reliability

Metropolis
Apply
5 hours ago

Base Salary

$204k - $249k/yr

Responsibilities

  • Own the reliability posture, metrics, and systems supporting 99.9%+ uptime across the Metropolis platform.
  • Design automatic failover for external dependencies including Twilio and Stripe using circuit breakers, retry policies, and degraded-mode operations.
  • Build multi-region deployment strategies with database replication, automated failover, DNS-based traffic routing, and disaster recovery planning and testing.
  • Establish Datadog monitoring for APM, logs, metrics correlation, synthetic monitoring, SLO-based alerting, on-call rotation, escalation policies, and service health dashboards.
  • Own incident management workflows, tooling, post-mortems, runbook automation, and MTTR reduction initiatives.
  • Drive adoption of health checks, graceful degradation, feature flags, rate limiting, backpressure, and chaos engineering practices.
  • Build local dependency mirrors with artifact caching, dependency pinning, and vulnerability scanning.

Requirements

  • 8+ years of engineering experience spanning software engineering, reliability engineering, SRE practices, or production operations at scale.
  • Expertise with multi-region architectures, failover automation, circuit breakers, chaos engineering, and disaster recovery.
  • Deep production observability experience with monitoring, alerting, tracing, and logging systems at scale, specifically Datadog or similar APM platforms.
  • Experience designing resilient distributed systems that handle failures, network partitions, and external dependency outages.
  • Knowledge of database replication, backup and restore, connection pooling, query optimization, and relational and NoSQL databases.
  • Production AWS experience, including multi-region deployments, load balancing, and DNS-based failover.
  • Experience with AI-powered development tools such as Claude Code, GitHub Copilot, or similar agentic coding tools, particularly context engineering.
  • Strong technical communication skills and experience influencing decisions, documenting systems, conducting post-mortems, and establishing reliability standards.
  • Expert-level Java and/or Scala proficiency, including JVM performance, concurrency, and operational characteristics.
  • Preferred qualifications include Scala experience, SRE or reliability engineering experience at operationally excellent companies, incident response leadership, chaos engineering with tools such as Chaos Monkey or Gremlin, hyperscale performance optimization, and relevant open-source contributions or technical writing.

Benefits

  • Office-first work arrangement requiring employees to be onsite at least four days per week.
  • Base salary range of $203,500 to $248,500 USD annually, plus potential healthcare benefits, 401(k), disability coverage, life insurance, stock options, and bonus plans.

Tech Stack

AWSDatadogGitJavaMySQLPostgreSQLReactScalaSnowflakeTypeScript

Categories

Site Reliability
Contact me