13 days ago
Hyderābād, IndiaSenior
Responsibilities
- Monitor system health, performance metrics, and alerts and respond promptly to incidents.
- Diagnose issues, troubleshoot problems, and help restore services in collaboration with senior SREs and other teams.
- Assist with software application deployments, release activities, and infrastructure changes while minimizing downtime.
- Automate routine tasks and improve operational efficiency with SRE and operations teams.
- Support capacity planning, resource-utilization monitoring, infrastructure scaling recommendations, and reliability improvements.
- Document incidents and resolutions, participate in post-incident reviews, contribute to root cause analysis, and implement preventive measures.
- Collaborate with security teams on security best practices, compliance, and security incident response.
- Stay current with Site Reliability Engineering trends, technologies, and best practices while developing technical skills.
Requirements
- Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
- Moderate hands-on experience in Site Reliability Engineering or related roles, including designing and maintaining highly available and scalable systems.
- Moderate experience with incident response procedures and troubleshooting techniques.
- Moderate experience with automation principles and tools such as Terraform, Jenkins, and Git.
- Developing knowledge of scripting or programming languages such as Python, Bash, or PowerShell.
- Familiarity with cloud platforms such as AWS, Azure, and Google Cloud, networking, system administration, Linux/Unix systems, and command-line tools.
- Developing expertise with performance monitoring, optimization, and troubleshooting tools such as Prometheus, Grafana, or New Relic.
- Understanding of security principles, compliance requirements, incident management, monitoring, and configuration management.
- Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator are preferred.
- Strong problem-solving, analytical, communication, collaboration, and attention-to-detail skills.
Benefits
- On-site work arrangement.
- Opportunities for training, certifications, self-study, and professional growth.
- Work with experienced SRE professionals and cross-functional development, operations, and security teams.
