
Cloud Site Reliability Engineer II
Barracuda Networks, Inc.4 days ago
Ann Arbor, MI, USAMid Level
Responsibilities
- Operate, scale, and automate centralized Loki, Mimir/Prometheus, Tempo, and Grafana telemetry infrastructure.
- Design Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
- Establish alerting strategies, SLO/SLI tracking, and notification routing to identify production degradation proactively.
- Automate log collector, metric exporter, and monitoring-agent deployments across EKS and AKS using ArgoCD and Terragrunt.
- Contribute to Kubernetes platform health, performance tuning, and infrastructure modernization.
- Build automation, diagnostic tooling, and operational scripts using AI-assisted developer tools.
- Collaborate with product and tenant teams on observability onboarding, distributed tracing instrumentation, and performance troubleshooting.
Requirements
- 2–4 years of experience working with public cloud infrastructure, including AWS and/or Azure, with a focus on observability, monitoring, and systems reliability.
- 1–2+ years of hands-on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes, including EKS or AKS.
- Practical experience configuring, operating, or building dashboards with modern observability stacks such as Grafana, ELK, or Splunk.
- Working knowledge of Terraform and/or Terragrunt and GitOps workflows using ArgoCD or Flux.
- Solid Python or Bash scripting skills for system automation, telemetry pipelines, and operational tooling; Go is a plus.
- Eagerness to use AI coding agents such as Claude Code, OpenCode, Codex CLI, and GitHub Copilot.
- Strong analytical troubleshooting, communication, and collaboration skills.
Benefits
- Equity in the form of non-qualifying options.
- High-quality health benefits.
- Retirement plan with employer match.
- Career-growth, cross-training, and internal mobility opportunities.
- Flexible Time Off and Paid Time Off benefits.
- Volunteer opportunities.
- Hybrid work arrangement, indicated by #LI-hybrid.
Categories
DevOpsSite Reliability
About Barracuda Networks, Inc.
Barracuda Networks builds cybersecurity products and services for businesses and managed service providers, spanning email protection, application and network security, data protection/backup, and managed XDR delivered as appliances and SaaS. Founded in 2003 and headquartered in Campbell, CA, it is privately held by Kohlberg Kravis Roberts & Co. Customers use Barracuda to secure cloud and on‑premise environments across Microsoft 365 and other enterprise workloads.