Scopely

Senior Platform Engineer (Reliability) - Unannounced Project

Scopely
Apply
22 days ago
Dublin, Ireland +2 moreSenior
H1B Sponsor

Responsibilities

  • Operate, monitor, and continuously improve NATS and NATS JetStream messaging infrastructure.
  • Design and own metrics, logs, and traces strategies for distributed backend systems.
  • Correlate infrastructure and application signals to diagnose distributed-system root causes.
  • Define, implement, and maintain SLO frameworks and error budgets for backend services and messaging pipelines.
  • Build reliability tooling, runbooks, and operational frameworks as internal products.
  • Partner with backend, infrastructure, and product engineers on reliability standards and architecture decisions.
  • Lead or contribute to incident response and drive postmortems toward systemic fixes.
  • Review infrastructure and backend code to identify instrumentation gaps, operational anti-patterns, inadequate error handling, and operational risks.

Requirements

  • Strong background in Site Reliability Engineering, production operations, or backend engineering with significant operational focus.
  • Production experience with NATS or NATS JetStream, or equivalent distributed messaging systems such as Apache Kafka or AWS Kinesis.
  • Experience with leader election, RAFT consensus, and log replication.
  • Ability to navigate C#, Go, Python, or similar application codebases and identify code-level operational issues.
  • Strong experience designing and implementing observability strategies for distributed systems.
  • Experience debugging complex distributed systems across infrastructure and application layers.
  • Solid understanding of AWS and containerized workloads such as ECS and Kubernetes/EKS.
  • Ability to communicate operational risks across technical and non-technical stakeholders.
  • Advocacy for managing infrastructure, SLOs, monitors, and dashboards through Infrastructure as Code.
  • Experience with NATS JetStream stream configuration, consumer groups, retention policies, and delivery guarantees is preferred.
  • Experience with Datadog or equivalent observability tooling is preferred.
  • Exposure to database query performance, connection pooling, replication lag, and data modeling is preferred.
  • Experience mentoring engineers or driving reliability culture across teams is preferred.

Benefits

  • Remote or hybrid work is available in Spain, Ireland, Portugal, or the UK.
  • Visa sponsorship and relocation assistance are available from any location.
  • Opportunity to work on an ambitious unannounced multiplayer strategy/MMO game.

Categories

DevOpsSite Reliability
Scopely

About Scopely

1,001-5,000 employees

Scopely is a leading global interactive entertainment and video game company, home to many top-grossing, award-winning franchises, including the most successful mobile game ever launched "MONOPOLY GO!," along with "Stumble Guys," "Star Trek™ Fleet Command," "MARVEL Strike Force," "WWE Champions" and "Yahtzee® With Buddies," among many others. Scopely creates, publishes, and live-operates immersive games across mobile, web, PC and console that empower a directed-by-consumer™ experience. Founded in 2011, Scopely is fueled by a world-class team and a proprietary technology platform Playgami™ that supports one of the most diversified portfolios in the games industry. Recognized multiple times as one of Fast Company’s “World’s Most Innovative Companies” and as a TIME100 “Most Influential Company In the World,” Scopely has surpassed $10 billion in lifetime revenue due to its ability to create long-lasting game experiences that players enjoy for years. Scopely has global operations across North America, Central America, EMEA and Asia, with additional game studio partners across four continents. Scopely was acquired by Savvy Games Group in July 2023 for $4.9 billion, and is now an independent subsidiary of Savvy. For more information on Scopely, visit: scopely.com.

Contact me