Barracuda Networks, Inc.

Cloud Site Reliability Engineer II

Barracuda Networks, Inc.
Apply
4 days ago
Ottawa, CanadaMid Level

Responsibilities

  • Operate, scale, and automate centralized LGTM telemetry infrastructure using Loki, Mimir, Prometheus, Tempo, and Grafana.
  • Design Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
  • Establish alerting strategies, SLO/SLI tracking with Sloth, and notification routing to detect service degradation.
  • Automate deployment of log collectors, metric exporters, and monitoring agents across EKS and AKS using ArgoCD and Terragrunt.
  • Contribute to Kubernetes platform health, performance tuning, and infrastructure modernization.
  • Use Claude Code, OpenCode, and Codex CLI to build automation, diagnostic tooling, and operational scripts.
  • Collaborate with internal product and tenant teams on observability onboarding, distributed tracing instrumentation, and performance troubleshooting.

Requirements

  • 2–4 years of experience working with public cloud infrastructure, including AWS and/or Azure, with a focus on observability and systems reliability.
  • 1–2+ years of hands-on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes, including EKS or AKS.
  • Practical experience configuring, operating, or building dashboards with observability stacks such as Grafana, ELK, or Splunk.
  • Working knowledge of Terraform and/or Terragrunt and GitOps workflows using ArgoCD or Flux.
  • Solid Python or Bash scripting skills for system automation, telemetry pipelines, and operational tooling; Go is a plus.
  • Interest in using AI coding agents such as Claude Code, OpenCode, Codex CLI, and GitHub Copilot.
  • Strong analytical troubleshooting, communication, and collaboration skills.

Benefits

  • Equity in the form of non-qualifying options.
  • High-quality health benefits.
  • Retirement plan with employer match.
  • Career-growth opportunities and internal mobility, including cross-training.
  • Flexible Time Off and Paid Time Off benefits.
  • Volunteer opportunities.
  • Hybrid role located in Ottawa, Ontario.

Tech Stack

AWSAzureBashGitHub ActionsGoGrafanaHelmKubernetesPrometheusPythonSplunkTerraform

Categories

Site Reliability
Barracuda Networks, Inc.

About Barracuda Networks, Inc.

1,001-5,000 employees

Barracuda Networks builds cybersecurity products and services for businesses and managed service providers, spanning email protection, application and network security, data protection/backup, and managed XDR delivered as appliances and SaaS. Founded in 2003 and headquartered in Campbell, CA, it is privately held by Kohlberg Kravis Roberts & Co. Customers use Barracuda to secure cloud and on‑premise environments across Microsoft 365 and other enterprise workloads.

Contact me