GrepJob
Metropolis

Senior Staff Software Engineer, Reliability

Metropolis
Apply
about 2 hours ago
Bengaluru, IndiaSenior / Staff+
H1B Sponsor

Responsibilities

  • Own the overall reliability posture for the Metropolis platform.
  • Design and implement automatic failover mechanisms for critical external dependencies.
  • Architect and build regional deployment strategies with database replication and automated failover.
  • Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation.
  • Implement synthetic monitoring, SLO-based alerting, and service health dashboards.
  • Own the incident management process including workflows and tooling.
  • Drive adoption of resilience patterns across all services.
  • Build and maintain local mirrors for critical dependencies.

Requirements

  • 8+ years of engineering experience in software engineering or reliability engineering.
  • Expert-level reliability engineering skills with experience in multi-region architectures.
  • Hands-on experience with monitoring, alerting, and logging systems at scale.
  • Strong systems thinking with proven ability to design resilient distributed systems.
  • Knowledge of database replication strategies and data systems.
  • Experience operating systems on AWS with multi-region deployments.
  • Familiarity with AI-powered development tools for enhanced productivity.
  • Proficiency in Java and/or Scala with understanding of JVM performance.

Benefits

  • In-person collaboration to drive innovation and strengthen culture.
  • Office-first model requiring employees to be on-site at least four days a week.

Tech Stack

AWSDatadogGitJavaMySQLPostgreSQLReactScalaSnowflakeTypeScript

Categories

AI & MLBackendData EngineeringDevOps