TCGplayer

Site Reliability Engineer 3

TCGplayer
Apply
2 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Monitor the health of critical eBay services and address potential issues proactively.
  • Collaborate with Architecture, Engineering, and Operations teams to develop solutions for high availability, reliability, and performance.
  • Resolve recurring technical issues, onboard alerts, and develop Standard Operating Procedures.
  • Build and improve monitoring and incident-mitigation tools.
  • Conduct reliability audits and tests to strengthen incident-management capabilities.
  • Act as Incident Commander during major incidents, manage alarms, and communicate with leadership and partner teams.

Requirements

  • 4+ years of professional software engineering experience, ideally on backend or platform teams.
  • Proficiency in one or more programming languages such as Java, Go, or Python.
  • Strong incident-management, leadership, technical-triage, and troubleshooting skills, particularly during crises.
  • Familiarity with cloud platforms, Kubernetes, and infrastructure-as-code tools.
  • Experience with observability stacks such as Prometheus, Grafana, ELK, or OpenTelemetry.
  • Strong interpersonal and communication skills for fast-paced, dynamic environments.

Tech Stack

GoGrafanaJavaKubernetesPrometheusPython

Categories

Site Reliability
TCGplayer

About TCGplayer

201-500 employees
Contact me