Xsolla

Operations Engineer, Germany

Xsolla
Apply
2 months ago
Berlin, GermanySenior

Responsibilities

  • Monitor Datadog dashboards and correlate APM, logs, metrics, synthetic tests, and Real User Monitoring signals to identify anomalies and determine incident significance.
  • Create and triage production incident tickets in JIRA Service Management, investigate technical symptoms, assess blast radius, identify likely root-cause domains, and route incidents to the appropriate team.
  • Own lower-severity incidents from detection through resolution, following runbooks and escalating issues that exceed defined thresholds or require code-level fixes.
  • Support major incident war rooms by providing real-time operational data, maintaining incident timelines and evidence, and executing mitigation actions.
  • Draft Slack updates, stakeholder notifications, customer-facing status-page communications, incident timelines, and initial post-incident review documents.
  • Analyze incident trends, recurring issues, and production bugs using Datadog, JIRA, and Slack, and contribute findings to operational reports.
  • Build and maintain alert-enrichment scripts, incident templates, Slack workflows, dashboard widgets, and operational runbooks.
  • Conduct shift handoffs, participate in SRE knowledge-transfer sessions, publish periodic critical-application health reports, and cover basic Incident Commander duties when needed.

Requirements

  • Previous experience working at a gaming company is required.
  • At least 4 years of experience in SRE, DevOps, production operations, NOC, or technical operations supporting high-availability environments.
  • Strong troubleshooting and investigation skills across application logs, APM traces, infrastructure metrics, database queries, and network paths.
  • Hands-on Datadog or equivalent observability-platform experience, including APM, log queries, infrastructure dashboards, SLO burn rates, monitors, and alerts.
  • Proficiency in at least one scripting language: Python, Go, or Bash.
  • Clear written and verbal English communication skills for incident tickets, investigation notes, Slack updates, shift handoffs, status-page communications, and post-incident review drafts.
  • Working knowledge of Kubernetes and cloud infrastructure, with GCP preferred and AWS or Azure acceptable.
  • Understanding of SLOs, error budgets, and multi-window burn-rate alerting.
  • Experience with JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.
  • Experience with or strong interest in AI/ML-assisted operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
  • Availability for 24x7 follow-the-sun shift operations and rotating weekend on-call.
  • Familiarity with Datadog Service Catalog, synthetic monitoring, RUM, distributed-systems debugging, database operations, CI/CD and deployment tooling, or JIRA Service Management administration is preferred.
  • ITIL Foundation certification is a plus but not required.

Benefits

  • Unlimited Flexible Time Off.
  • Gym membership and monthly train ticket.
  • Personalized career roadmap, training, and educational opportunities.
  • 24x7 follow-the-sun shift arrangement with handoff overlaps and rotating weekend on-call.
  • Background checks may be conducted after the final interview where permitted by law.

Tech Stack

Apache KafkaAWSAzureBashDatadogGitLab CI/CDGoGoogle Cloud PlatformGrafanaHelmKubernetesMySQLPostgreSQLPythonRedisSplunk

Categories

Site Reliability
Xsolla

About Xsolla

1,001-5,000 employees
Contact me