2 months ago
Berlin, GermanySenior
Responsibilities
- Monitor Datadog dashboards and correlate APM, logs, metrics, synthetic tests, and Real User Monitoring signals to identify anomalies and determine incident significance.
- Create and triage production incident tickets in JIRA Service Management, investigate technical symptoms, assess blast radius, identify likely root-cause domains, and route incidents to the appropriate team.
- Own lower-severity incidents from detection through resolution, following runbooks and escalating issues that exceed defined thresholds or require code-level fixes.
- Support major incident war rooms by providing real-time operational data, maintaining incident timelines and evidence, and executing mitigation actions.
- Draft Slack updates, stakeholder notifications, customer-facing status-page communications, incident timelines, and initial post-incident review documents.
- Analyze incident trends, recurring issues, and production bugs using Datadog, JIRA, and Slack, and contribute findings to operational reports.
- Build and maintain alert-enrichment scripts, incident templates, Slack workflows, dashboard widgets, and operational runbooks.
- Conduct shift handoffs, participate in SRE knowledge-transfer sessions, publish periodic critical-application health reports, and cover basic Incident Commander duties when needed.
Requirements
- Previous experience working at a gaming company is required.
- At least 4 years of experience in SRE, DevOps, production operations, NOC, or technical operations supporting high-availability environments.
- Strong troubleshooting and investigation skills across application logs, APM traces, infrastructure metrics, database queries, and network paths.
- Hands-on Datadog or equivalent observability-platform experience, including APM, log queries, infrastructure dashboards, SLO burn rates, monitors, and alerts.
- Proficiency in at least one scripting language: Python, Go, or Bash.
- Clear written and verbal English communication skills for incident tickets, investigation notes, Slack updates, shift handoffs, status-page communications, and post-incident review drafts.
- Working knowledge of Kubernetes and cloud infrastructure, with GCP preferred and AWS or Azure acceptable.
- Understanding of SLOs, error budgets, and multi-window burn-rate alerting.
- Experience with JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.
- Experience with or strong interest in AI/ML-assisted operations such as anomaly detection, alert correlation, predictive monitoring, or automated remediation.
- Availability for 24x7 follow-the-sun shift operations and rotating weekend on-call.
- Familiarity with Datadog Service Catalog, synthetic monitoring, RUM, distributed-systems debugging, database operations, CI/CD and deployment tooling, or JIRA Service Management administration is preferred.
- ITIL Foundation certification is a plus but not required.
Benefits
- Unlimited Flexible Time Off.
- Gym membership and monthly train ticket.
- Personalized career roadmap, training, and educational opportunities.
- 24x7 follow-the-sun shift arrangement with handoff overlaps and rotating weekend on-call.
- Background checks may be conducted after the final interview where permitted by law.
Tech Stack
Apache KafkaAWSAzureBashDatadogGitLab CI/CDGoGoogle Cloud PlatformGrafanaHelmKubernetesMySQLPostgreSQLPythonRedisSplunk
Categories
Site Reliability
