Barracuda Networks, Inc.

Cloud Site Reliability Senior Engineer

Barracuda Networks, Inc.
Apply
2 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Improve the reliability, availability, and operational excellence of production services through shared team ownership.
  • Lead and contribute to troubleshooting, incident response, root-cause analysis, and preventive improvements across cloud services and supporting platforms.
  • Partner with Engineering, Cloud Operations, and SRE teams on release readiness, deployment practices, and operational standards.
  • Support cloud platform improvements, automation, resilience initiatives, and cost-conscious operations across Azure, AWS, Terraform-managed infrastructure, and containerized services.
  • Strengthen observability, alerting, and operational insights while contributing to incident response and service health improvements.
  • Participate in on-call rotation if required.
  • Share knowledge, improve documentation, and reduce product and domain silos across SRE and Cloud Operations teams.

Requirements

  • Bachelor’s degree in computer science engineering, Information Technology, or an equivalent degree.
  • 5+ years of progressive experience in Site Reliability Engineering, Cloud Operations, DevOps, or Platform Engineering.
  • Strong Linux/Unix command-line administration, troubleshooting, package management, and systems operations skills.
  • Experience managing production cloud infrastructure across Azure and/or AWS.
  • Hands-on experience operating container orchestration platforms, including AKS migrations, upgrades, and production operations.
  • Hands-on Infrastructure as Code experience, preferably with Terraform.
  • Automation and scripting experience using Python, Bash, Go, YAML, generative AI tools, or similar technologies.
  • Experience with CI/CD pipelines using Azure DevOps, ArgoCD, Jenkins, GitHub Actions, or similar platforms.
  • Experience with Docker and container registries such as Azure Container Registry.
  • Experience with configuration management and automation tools such as Ansible, Puppet, or Chef.
  • Experience with monitoring, observability, and incident response platforms such as PagerDuty, Grafana, Prometheus, ELK/OpenSearch, New Relic, or Sensu.
  • Understanding of networking, the OSI model, SQL/NoSQL databases, and distributed-systems troubleshooting.
  • Ability to prioritize tasks, work independently, and communicate clearly with technical and non-technical audiences.
  • Preferred experience includes large-scale AKS migrations, container platform modernization, cloud transformation, complex production release coordination, AI-assisted operations, reliability tooling, self-healing capabilities, automation frameworks, and relevant Azure, AWS, Terraform, container-platform, or cloud-reliability certifications.

Benefits

  • Internal mobility and cross-training opportunities are available, with support for defining future career steps within Barracuda.
  • The company promotes an inclusive environment where employees can voice opinions, make an impact, and explore areas of interest.
  • The role may involve participation in an on-call rotation.
  • The position involves collaboration with globally distributed teams through video conferencing, Slack, and other communication tools.

Tech Stack

AnsibleAWSAzureBashChefDockerGitHub ActionsGoGrafanaJenkinsLinuxPrometheusPuppetPythonSQLTerraform

Categories

DevOpsSite Reliability
Barracuda Networks, Inc.

About Barracuda Networks, Inc.

1,001-5,000 employees
Contact me