2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Monitor the health of critical eBay services and address potential issues proactively.
- Collaborate with Architecture, Engineering, and Operations teams to develop solutions for high availability, reliability, and performance.
- Resolve recurring technical issues, onboard alerts, and develop Standard Operating Procedures.
- Build and improve monitoring and incident-mitigation tools.
- Conduct reliability audits and tests to strengthen incident-management capabilities.
- Act as Incident Commander during major incidents, manage alarms, and communicate with leadership and partner teams.
Requirements
- 4+ years of professional software engineering experience, ideally on backend or platform teams.
- Proficiency in one or more programming languages such as Java, Go, or Python.
- Strong incident-management, leadership, technical-triage, and troubleshooting skills, particularly during crises.
- Familiarity with cloud platforms, Kubernetes, and infrastructure-as-code tools.
- Experience with observability stacks such as Prometheus, Grafana, ELK, or OpenTelemetry.
- Strong interpersonal and communication skills for fast-paced, dynamic environments.
Tech Stack
Categories
Site Reliability
