1 month ago
Milan, ItalyMid Level
Responsibilities
- Build and maintain core infrastructure as code using Terraform and Ansible.
- Implement monitoring, logging, and alerting systems to support 99.99% uptime.
- Debug and resolve production incidents and lead blameless post-mortems.
- Write automation to reduce operational toil and enable self-service for engineering teams.
- Collaborate with developers on reliability and scalability practices.
- Contribute to capacity planning, disaster recovery drills, and security hardening.
- Participate in a sustainable on-call rotation.
Requirements
- Experience operating production workloads on a major cloud provider such as AWS, GCP, or Azure.
- Proficiency in at least one programming or scripting language such as Golang, Python, or Bash.
- Hands-on experience with Docker and Kubernetes.
- Knowledge of infrastructure as code principles and tools; Terraform experience is a plus.
- Familiarity with CI/CD concepts and pipeline tools such as GitLab CI or Jenkins.
- Understanding of observability stacks such as Prometheus, Grafana, or ELK.
Benefits
- Participation in a fair and sustainable on-call rotation.
- Opportunity to work on large-scale infrastructure supporting critical customer applications.
Tech Stack
AnsibleAWSAzureBashDockerGitLab CI/CDGoGoogle Cloud PlatformGrafanaJenkinsKubernetesPrometheusPythonTerraform
Categories
Site Reliability
