Senior Staff Software Engineer, Reliability
Metropolisabout 2 hours ago
Bengaluru, IndiaSenior / Staff+
H1B Sponsor
Responsibilities
- Own the overall reliability posture for the Metropolis platform.
- Design and implement automatic failover mechanisms for critical external dependencies.
- Architect and build regional deployment strategies with database replication and automated failover.
- Establish comprehensive monitoring using Datadog for APM, logs, and metrics correlation.
- Implement synthetic monitoring, SLO-based alerting, and service health dashboards.
- Own the incident management process including workflows and tooling.
- Drive adoption of resilience patterns across all services.
- Build and maintain local mirrors for critical dependencies.
Requirements
- 8+ years of engineering experience in software engineering or reliability engineering.
- Expert-level reliability engineering skills with experience in multi-region architectures.
- Hands-on experience with monitoring, alerting, and logging systems at scale.
- Strong systems thinking with proven ability to design resilient distributed systems.
- Knowledge of database replication strategies and data systems.
- Experience operating systems on AWS with multi-region deployments.
- Familiarity with AI-powered development tools for enhanced productivity.
- Proficiency in Java and/or Scala with understanding of JVM performance.
Benefits
- In-person collaboration to drive innovation and strengthen culture.
- Office-first model requiring employees to be on-site at least four days a week.