TCGplayer

MTS 1, Site Reliability Engineer

TCGplayer
Apply
16 hours ago
Bengaluru, IndiaEntry Level

Responsibilities

  • Monitor the health of critical services and address potential issues proactively.
  • Collaborate with architecture, engineering, and operations teams to improve service availability, reliability, and performance.
  • Resolve recurring technical issues, onboard new alerts, and develop Standard Operating Procedures.
  • Build and enhance monitoring and incident-mitigation tools and conduct reliability audits and tests.
  • Act as Incident Commander during major incidents, manage alarms, and communicate with leadership and partner teams.

Requirements

  • 6+ years of professional software engineering experience, ideally with backend or platform teams.
  • Proficiency in one or more programming languages such as Java, Go, or Python.
  • Strong incident management, technical triage, troubleshooting, and crisis leadership skills.
  • Familiarity with cloud platforms, Kubernetes or other container orchestration, and infrastructure-as-code tools.
  • Experience with observability stacks such as Prometheus, Grafana, ELK, or OpenTelemetry.
  • Strong interpersonal and communication skills.

Tech Stack

GoGrafanaJavaKubernetesPrometheusPython

Categories

Site Reliability
TCGplayer

About TCGplayer

201-500 employees

TCGplayer operates an online marketplace and software tools for buying and selling trading card games and related collectibles, serving hobby shops and individual sellers. It generates revenue through marketplace fees, seller subscriptions, and fulfillment programs such as TCGplayer Direct. An eBay subsidiary, the company provides authentication and logistics services that help stores list inventory at scale and reach buyers across the U.S. and internationally.

Contact me