6 months ago
Athens, GreeceMid Level
Responsibilities
- Serve as the on-call first responder for platform incidents and lead end-to-end triage and resolution.
- Coordinate with DevOps and Engineering and own post-incident reviews and corrective actions.
- Own alerting, dashboards, log aggregation, distributed tracing, and service health standards.
- Define and track Service Level Indicators and Objectives, analyze reliability trends, and reduce alert noise.
- Drive platform, configuration, architecture, and operational improvements to reduce incident frequency and blast radius.
- Build runbooks, self-healing scripts, and automated remediation tooling.
- Partner with Engineers and DevOps on reliability in releases, capacity planning, load testing, and infrastructure hardening.
- Participate in release planning as a reliability gate for critical deployments.
Requirements
- At least 3 years of experience in SRE, platform engineering, or a senior TechOps role.
- Hands-on experience owning an on-call rotation.
- A track record of leading post-incident reviews and delivering action items.
- Experience in iGaming, fintech, or another regulated 24/7 industry is a significant advantage.
- Ability to make sound decisions under pressure and investigate the underlying causes of failures.
Benefits
- Hybrid workplace in Athens, Greece.
- Full-time employment.
- Competitive salary and bonus scheme.
- Group health and medical insurance package.
- Equipment provided for the role.
- Career development, performance management, and training opportunities.
- Free access to the in-house gym.
- Shuttle buses and carpooling options.
- International, inclusive, and multicultural team environment.
- Company events, sports, and team-building activities.
Categories
Site Reliability
