5 months ago
Milan, ItalyMid Level
Responsibilities
- Provide operational support for production environments and maintain service availability and reliability.
- Develop and maintain automation scripts and tools using Bash, Python, and Ansible.
- Monitor system performance, identify issues, and implement preventative solutions.
- Participate in on-call rotation, respond to incidents, perform root cause analysis, and drive resolution.
- Improve configuration management, deployment practices, CI/CD pipelines, and infrastructure as code.
- Document operational processes, troubleshooting procedures, and automation workflows.
- Collaborate with development, QA, and infrastructure teams on operational best practices.
Requirements
- Proven experience in production support or Site Reliability Engineering within a complex, high-availability environment.
- Strong proficiency with Bash, Python, and Ansible for automation.
- Experience with monitoring and alerting tools such as Prometheus, Grafana, Elastic Stack, or Datadog.
- Strong Linux/Unix systems administration and troubleshooting skills.
- Familiarity with AWS and containerization technologies such as Docker and Kubernetes.
- Experience with infrastructure as code or configuration management tools such as Terraform and CloudFormation.
- Knowledge of networking fundamentals, security best practices, and incident management processes.
- Strong problem-solving, communication, collaboration, and attention-to-detail skills, including the ability to work under pressure.
- Experience with Git, database administration or troubleshooting using MySQL, PostgreSQL, or Oracle, and automation languages such as Go or Ruby is desirable.
Tech Stack
Categories
Site Reliability