2 months ago
Remote, United StatesSenior
Responsibilities
- Own service reliability and operational health by establishing SLOs and SLIs and improving availability and performance across cloud platforms.
- Lead incident response coordination, troubleshooting, root cause analysis, post-incident reviews, remediation, and recovery activities.
- Design reliability-focused automation, operational tooling, runbooks, monitoring, logging, alerting, event correlation, and escalation procedures.
- Conduct performance and capacity analysis and recommend scaling, performance, capacity-planning, and operational-readiness improvements.
- Partner with engineering teams on deployment readiness, rollback readiness, production supportability, and operational best practices.
- Contribute to disaster recovery and business continuity planning and conduct operational readiness exercises.
- Mentor team members and establish reliability standards, operational documentation, standard operating procedures, and knowledge-base materials.
- Support cloud governance, compliance, security initiatives, access control, tagging, logging, and audit readiness.
Requirements
- Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field, or an equivalent combination of education and experience.
- 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, DevOps, Infrastructure Operations, or a related discipline with demonstrated production-system ownership.
- Hands-on experience supporting production cloud environments using Google Cloud Platform, AWS, or equivalent providers.
- Expertise in monitoring, observability, alerting, incident response, root cause analysis, and production support in distributed or cloud-native architectures.
- Experience with Infrastructure as Code and version-control best practices, including Terraform, Deployment Manager, or CloudFormation.
- Experience with incident management, post-incident reviews, corrective actions, Kubernetes operations, containerization, orchestration, application performance monitoring, and distributed tracing.
- Experience mentoring junior engineers or leading operational improvement initiatives.
- Required or equivalent certifications include Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, Google Cloud Professional Data Engineer, or AWS SysOps Administrator.
- Advanced certifications in Kubernetes, Terraform, observability platforms, DevOps, Site Reliability Engineering, or ITIL are preferred.
- Knowledge of cloud governance, compliance frameworks, disaster recovery, business continuity, networking, security, compute models, and observability platforms.
- Proficiency in scripting languages such as Python, Bash, or Go and ability to diagnose complex infrastructure issues and influence cross-functional teams.
Tech Stack
Categories
DevOpsSite Reliability
About NextGen
NEXTGEN is an Australia-based technology services and value-added distribution company that helps vendors and channel partners sell cybersecurity, cloud, enterprise software, and data management solutions. It offers software licensing, compliance and audit services, data centre and storage solutions, and go-to-market and digital marketing support. Founded in 2011 and headquartered in North Sydney, it is now part of Exclusive Networks, expanding its reach across international markets.
