4 days ago
Bengaluru, India or Hyderābād, IndiaSenior
Responsibilities
- Monitor system health, performance metrics, and alerts; troubleshoot issues and restore services during incidents.
- Implement incident response processes, lead response efforts, and conduct root cause and post-incident analyses.
- Design and maintain automation tools, scripts, frameworks, deployment processes, and self-healing capabilities.
- Implement infrastructure as code and enforce configuration consistency and standards across environments.
- Optimize system performance, scalability, resource usage, availability, and operational efficiency.
- Plan for capacity growth and forecast system resource needs.
- Collaborate with development, operations, security, and other stakeholders on reliability goals and knowledge sharing.
- Implement security best practices, assess vulnerabilities, and support compliance with security standards and regulations.
Requirements
- Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
- Seasoned hands-on experience in Site Reliability Engineering or related roles designing and maintaining highly available, scalable systems.
- Seasoned expertise with Linux/Unix systems, networking, system administration, cloud platforms, and associated services.
- Proficiency in multiple programming or scripting languages such as Python, Java, Go, Ruby, Bash, or PowerShell.
- Experience with infrastructure architecture, infrastructure-as-code tools, containerization, automation frameworks, CI/CD pipelines, and deployment strategies.
- Experience with incident management, complex troubleshooting, post-incident analysis, root cause analysis, and incident response leadership.
- Experience with performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
- Understanding of security principles, security controls, security assessments, compliance requirements, DevOps principles, and Agile methodologies.
- Strong problem-solving, analytical, communication, collaboration, leadership, and continuous-improvement skills.
- Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator are preferred.
Benefits
- Hybrid working arrangement.
