NTT

Senior Associate Site Reliability Engineer

NTT
Apply
6 days ago
Hyderābād, IndiaSenior

Responsibilities

  • Monitor system health, performance metrics, and alerts and respond promptly to incidents.
  • Diagnose issues, troubleshoot problems, and help restore services with senior SREs and cross-functional teams.
  • Assist with application and infrastructure deployments, releases, and practices that minimize downtime.
  • Automate routine operational tasks and improve efficiency with SRE and operations teams.
  • Support capacity planning, resource monitoring, infrastructure scaling recommendations, and reliability improvements.
  • Document incidents and resolutions, participate in post-incident reviews, and contribute to root cause analysis and preventive measures.
  • Collaborate with security teams to implement security best practices, compliance measures, and security incident mitigations.
  • Continue developing technical knowledge through training, certifications, industry research, and self-study.

Requirements

  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Moderate hands-on experience in Site Reliability Engineering or related roles, including designing and maintaining highly available and scalable systems.
  • Moderate experience with incident response procedures and troubleshooting techniques.
  • Moderate experience with automation principles and tools such as Terraform, Jenkins, and Git.
  • Developing knowledge of Python, Bash, or PowerShell; Linux/Unix systems; command-line tools; networking; system administration; and cloud platforms.
  • Developing expertise with performance monitoring, optimization, and troubleshooting tools such as Prometheus, Grafana, or New Relic.
  • Developing understanding of security principles, best practices, and compliance requirements.
  • Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator are preferred.
  • Strong problem-solving, analytical, communication, collaboration, and continuous-improvement skills.

Benefits

  • On-site working arrangement.
  • Opportunities for training, certifications, hands-on experience, and professional growth.

Tech Stack

AWSAzureBashGitGoogle CloudGrafanaJenkinsKubernetesLinuxPowerShellPrometheusPythonTerraform

Categories

Site Reliability
Contact me