
Site Reliability Engineer
Betsson Group1 day ago
Budapest, HungarySenior
Responsibilities
- Investigate system incidents, drive root cause analyses, and execute long-term remedial fixes.
- Proactively reduce incidents caused by system changes and lead blameless post-mortems.
- Define and enforce SLAs, SLOs, and success metrics for new initiatives.
- Build and maintain dashboards and comprehensive observability capabilities.
- Identify and resolve performance bottlenecks, optimize infrastructure and code, and conduct capacity planning.
- Maintain platform availability and functionality for users across multiple environments and data centres.
- Oversee deployments and change management to ensure new code does not disrupt existing systems.
- Automate manual tasks and build reliability tools to reduce operational toil.
Requirements
- Deep experience building dashboards and tracking SLAs and SLOs with observability and monitoring tools.
- Strong scripting and coding skills, with .NET, Python, PowerShell, or Bash preferred.
- Experience provisioning and managing infrastructure with Terraform or Ansible.
- Solid understanding of AWS, GCP, or Azure cloud platforms.
- Hands-on experience managing and scaling distributed systems with Kubernetes and Docker.
- Familiarity with GitLab CI, GitHub Actions, TeamCity, or Octopus deployment pipelines.
- Strong analytical skills for root cause analysis, calm incident-response skills, and experience leading blameless post-mortems.
- Experience with AWS cloud infrastructure, CDNs, and systems operating across multiple data centres and environments.
- Experience with cloud application load balancers, preferably AWS Application Load Balancer.
- Experience supporting cloud DNS services such as AWS Route 53, GCP Cloud DNS, or Azure DNS.
- Experience with Microsoft SQL databases, PostgreSQL, and Couchbase is an asset.
Benefits
- Hybrid work model with three days in the office and two days from home.
- Fitness and wellness allowance.
- Company mobile phone for private use with 100 GB of data.
- Annual HUF devaluation compensation.
- Private health insurance.
- Career development opportunities.
- Technical and soft-skill training opportunities.
- Breakfast, fruit, lunch, and team-building events.
Tech Stack
AnsibleAWSAzureBashCouchbaseDockerGitHub ActionsGitLab CI/CDGoogle Cloud PlatformGrafanaKubernetesMicrosoft SQL Server.NETOctopus DeployPostgreSQLPowerShellPrometheusPythonSplunkTeamCityTerraform
Categories
Site Reliability