
Site Reliability Engineer
ACI Worldwide17 days ago
Norcross, GA, USASenior
Responsibilities
- Define and manage SLOs, SLIs, error budgets, and service reliability metrics.
- Build and operate Azure-based platforms, including AKS, networking, security, and monitoring.
- Implement GitOps workflows with FluxCD or ArgoCD for Kubernetes and infrastructure.
- Design and maintain CI/CD pipelines using GitHub Actions or Azure DevOps.
- Automate infrastructure with Terraform, Bicep, and ARM.
- Lead incident response, root cause analysis, and reliability improvements.
- Establish observability through logs, metrics, tracing, and alerting.
- Embed DevSecOps practices including policy-as-code, vulnerability scanning, and secrets management.
- Partner with development teams to improve deployment safety and production readiness.
Requirements
- At least 3 years of experience in SRE, DevOps, or platform engineering.
- Strong expertise with Microsoft Azure, including AKS, monitoring, networking, and security.
- Hands-on experience with GitOps and Kubernetes.
- Proficiency with GitHub Actions or Azure DevOps for CI/CD.
- Infrastructure-as-Code experience with Terraform, Bicep, or ARM.
- Scripting and automation experience with Python, PowerShell, or Bash.
- Experience with incident management and operational best practices.
- Applicants must be currently authorized to work full-time in the United States.
Benefits
- Opportunities for growth and career development.
- Competitive compensation and benefits package.
- Innovative and collaborative work environment.
- Role is located in Norcross, GA or Omaha, NE and is hybrid.
Tech Stack
Categories
Site Reliability