17 hours ago
Dublin, IrelandSenior
Responsibilities
- Monitor the health of eBay’s critical services and proactively identify potential issues.
- Collaborate with architecture, engineering, and operations teams to develop solutions for high availability, reliability, and performance.
- Resolve recurring technical issues, onboard new alerts, and create high-quality Standard Operating Procedures.
- Build and improve monitoring and incident-mitigation tools and conduct reliability audits and tests.
- Serve as Incident Commander for major incidents, manage alarms, and communicate effectively with leadership and partner teams.
Requirements
- Have 4+ years of professional software engineering experience, ideally on backend or platform teams.
- Be proficient in one or more programming languages such as Java, Go, or Python.
- Demonstrate strong incident management, technical triage, troubleshooting, and leadership skills, particularly during crises.
- Have familiarity with cloud platforms, Kubernetes, and infrastructure-as-code tools.
- Have experience with observability stacks such as Prometheus, Grafana, ELK, or OpenTelemetry.
- Possess strong interpersonal and communication skills for fast-paced, dynamic environments.
Tech Stack
Categories
Site Reliability
About TCGplayer
TCGplayer operates an online marketplace and software tools for buying and selling trading card games and related collectibles, serving hobby shops and individual sellers. It generates revenue through marketplace fees, seller subscriptions, and fulfillment programs such as TCGplayer Direct. An eBay subsidiary, the company provides authentication and logistics services that help stores list inventory at scale and reach buyers across the U.S. and internationally.
