9 hours ago
Hyderābād, IndiaSenior
Responsibilities
- Monitor system health, performance metrics, and alerts; troubleshoot issues and restore services promptly.
- Lead incident response, root-cause analysis, post-incident reviews, and preventive remediation.
- Design and maintain automation tools, scripts, self-healing frameworks, deployment processes, and infrastructure-as-code implementations.
- Optimize system performance, scalability, capacity, resource usage, and configuration consistency.
- Implement monitoring, security, compliance, and operational reliability best practices in collaboration with relevant teams.
Requirements
- Seasoned hands-on experience in Site Reliability Engineering or related roles designing and maintaining highly available and scalable systems.
- Seasoned expertise with Linux/Unix systems, networking, system administration, cloud platforms, and associated services.
- Proficiency in multiple programming or scripting languages, including Python, Go, Java, Ruby, Bash, or PowerShell.
- Experience with infrastructure architectures, Terraform or CloudFormation, Docker or Kubernetes, automation frameworks, CI/CD pipelines, and deployment strategies.
- Experience with incident management, complex troubleshooting, root-cause analysis, post-incident reviews, and leading incident response.
- Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
- AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator certification is preferred.
Benefits
- On-site work arrangement.
