NTT

Senior Associate Site Reliability Engineer

NTT
Apply
13 days ago
Hyderābād, IndiaSenior

Responsibilities

  • Monitor system health, performance metrics, and alerts and respond promptly to incidents.
  • Diagnose issues, troubleshoot problems, and help restore services in collaboration with senior SREs and other teams.
  • Assist with software application deployments, release activities, and infrastructure changes while minimizing downtime.
  • Automate routine tasks and improve operational efficiency with SRE and operations teams.
  • Support capacity planning, resource-utilization monitoring, infrastructure scaling recommendations, and reliability improvements.
  • Document incidents and resolutions, participate in post-incident reviews, contribute to root cause analysis, and implement preventive measures.
  • Collaborate with security teams on security best practices, compliance, and security incident response.
  • Stay current with Site Reliability Engineering trends, technologies, and best practices while developing technical skills.

Requirements

  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Moderate hands-on experience in Site Reliability Engineering or related roles, including designing and maintaining highly available and scalable systems.
  • Moderate experience with incident response procedures and troubleshooting techniques.
  • Moderate experience with automation principles and tools such as Terraform, Jenkins, and Git.
  • Developing knowledge of scripting or programming languages such as Python, Bash, or PowerShell.
  • Familiarity with cloud platforms such as AWS, Azure, and Google Cloud, networking, system administration, Linux/Unix systems, and command-line tools.
  • Developing expertise with performance monitoring, optimization, and troubleshooting tools such as Prometheus, Grafana, or New Relic.
  • Understanding of security principles, compliance requirements, incident management, monitoring, and configuration management.
  • Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator are preferred.
  • Strong problem-solving, analytical, communication, collaboration, and attention-to-detail skills.

Benefits

  • On-site work arrangement.
  • Opportunities for training, certifications, self-study, and professional growth.
  • Work with experienced SRE professionals and cross-functional development, operations, and security teams.

Tech Stack

AWSAzureBashGitGoogle CloudGrafanaJenkinsKubernetesLinuxPowerShellPrometheusPythonTerraform

Categories

Site Reliability
Contact me