Senior Site Reliability Engineer
Resideo Technologies, Inc.2 months ago
Bucharest, RomaniaSenior
Responsibilities
- Design, implement, and optimize public cloud infrastructure across Azure, AWS, or GCP.
- Drive adoption and enhancement of Infrastructure-as-Code using Terraform, ARM Templates, or similar tools.
- Develop IT automation solutions with Ansible, Chef, and Helm for Kubernetes.
- Oversee CI/CD pipelines using Git, Git Actions, Jenkins, Docker, and Kubernetes.
- Architect and maintain observability and monitoring frameworks with Grafana, Prometheus, or Elastic.
- Manage critical incidents, conduct root cause analysis, and lead infrastructure upgrades to minimize downtime.
- Mentor engineering teams on cloud infrastructure, reliability engineering, and operational excellence.
Requirements
- At least 6 years of progressive experience in Site Reliability Engineering or a closely related cloud infrastructure role.
- At least 3 years of hands-on experience with a major public cloud platform and the ability to architect and manage cloud-native solutions.
- Experience designing and implementing Infrastructure-as-Code for complex environments, including at least 2 years with Terraform or similar platforms.
- Expertise with Docker, Kubernetes, and their supporting ecosystem.
- Advanced scripting experience with PowerShell, Bash, Python, or similar languages is valued.
- Experience with distributed systems, large-scale data platforms, or IoT infrastructure is valued.
- Experience providing strategic technical guidance and cross-functional leadership in a global team environment is valued.
- At least 5 years of experience administering and optimizing enterprise Windows and Linux environments is valued.
- Strong incident response, root cause analysis, and complex production problem-solving leadership is valued.
Benefits
- Flexible hybrid working arrangement.
- Meal ticket for each day worked.
- Medical coverage.
- 26 days of vacation.
Tech Stack
AnsibleAWSAzureBashChefDockerGitGoogle Cloud PlatformGrafanaHelmJenkinsKubernetesLinuxPowerShellPrometheusPythonTerraformWindows
Categories
DevOpsSite Reliability