
Cloud Site Reliability Senior Engineer
Barracuda Networks, Inc.2 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Improve the reliability, availability, and operational excellence of production services through shared team ownership.
- Lead and contribute to troubleshooting, incident response, root-cause analysis, and preventive improvements across cloud services and supporting platforms.
- Partner with Engineering, Cloud Operations, and SRE teams on release readiness, deployment practices, and operational standards.
- Support cloud platform improvements, automation, resilience initiatives, and cost-conscious operations across Azure, AWS, Terraform-managed infrastructure, and containerized services.
- Strengthen observability, alerting, and operational insights while contributing to incident response and service health improvements.
- Participate in on-call rotation if required.
- Share knowledge, improve documentation, and reduce product and domain silos across SRE and Cloud Operations teams.
Requirements
- Bachelor’s degree in computer science engineering, Information Technology, or an equivalent degree.
- 5+ years of progressive experience in Site Reliability Engineering, Cloud Operations, DevOps, or Platform Engineering.
- Strong Linux/Unix command-line administration, troubleshooting, package management, and systems operations skills.
- Experience managing production cloud infrastructure across Azure and/or AWS.
- Hands-on experience operating container orchestration platforms, including AKS migrations, upgrades, and production operations.
- Hands-on Infrastructure as Code experience, preferably with Terraform.
- Automation and scripting experience using Python, Bash, Go, YAML, generative AI tools, or similar technologies.
- Experience with CI/CD pipelines using Azure DevOps, ArgoCD, Jenkins, GitHub Actions, or similar platforms.
- Experience with Docker and container registries such as Azure Container Registry.
- Experience with configuration management and automation tools such as Ansible, Puppet, or Chef.
- Experience with monitoring, observability, and incident response platforms such as PagerDuty, Grafana, Prometheus, ELK/OpenSearch, New Relic, or Sensu.
- Understanding of networking, the OSI model, SQL/NoSQL databases, and distributed-systems troubleshooting.
- Ability to prioritize tasks, work independently, and communicate clearly with technical and non-technical audiences.
- Preferred experience includes large-scale AKS migrations, container platform modernization, cloud transformation, complex production release coordination, AI-assisted operations, reliability tooling, self-healing capabilities, automation frameworks, and relevant Azure, AWS, Terraform, container-platform, or cloud-reliability certifications.
Benefits
- Internal mobility and cross-training opportunities are available, with support for defining future career steps within Barracuda.
- The company promotes an inclusive environment where employees can voice opinions, make an impact, and explore areas of interest.
- The role may involve participation in an on-call rotation.
- The position involves collaboration with globally distributed teams through video conferencing, Slack, and other communication tools.
Tech Stack
Categories
DevOpsSite Reliability