NTT

Site Reliability Engineer

NTT
Apply
4 days ago
Bengaluru, India or Hyderābād, IndiaSenior

Responsibilities

  • Monitor system health, performance metrics, and alerts; troubleshoot issues and restore services during incidents.
  • Implement incident response processes, lead response efforts, and conduct root cause and post-incident analyses.
  • Design and maintain automation tools, scripts, frameworks, deployment processes, and self-healing capabilities.
  • Implement infrastructure as code and enforce configuration consistency and standards across environments.
  • Optimize system performance, scalability, resource usage, availability, and operational efficiency.
  • Plan for capacity growth and forecast system resource needs.
  • Collaborate with development, operations, security, and other stakeholders on reliability goals and knowledge sharing.
  • Implement security best practices, assess vulnerabilities, and support compliance with security standards and regulations.

Requirements

  • Bachelor's degree or equivalent in Computer Science, Information Technology, or a related field.
  • Seasoned hands-on experience in Site Reliability Engineering or related roles designing and maintaining highly available, scalable systems.
  • Seasoned expertise with Linux/Unix systems, networking, system administration, cloud platforms, and associated services.
  • Proficiency in multiple programming or scripting languages such as Python, Java, Go, Ruby, Bash, or PowerShell.
  • Experience with infrastructure architecture, infrastructure-as-code tools, containerization, automation frameworks, CI/CD pipelines, and deployment strategies.
  • Experience with incident management, complex troubleshooting, post-incident analysis, root cause analysis, and incident response leadership.
  • Experience with performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
  • Understanding of security principles, security controls, security assessments, compliance requirements, DevOps principles, and Agile methodologies.
  • Strong problem-solving, analytical, communication, collaboration, leadership, and continuous-improvement skills.
  • Relevant certifications such as AWS Certified DevOps Engineer - Professional, Google Cloud Professional DevOps Engineer, or Certified Kubernetes Administrator are preferred.

Benefits

  • Hybrid working arrangement.

Tech Stack

AWSAzureBashCircleCIDockerGitLab CI/CDGoGoogle CloudGrafanaJavaJenkinsKubernetesLinuxPowerShellPrometheusPythonRubyTerraform

Categories

Site Reliability
Contact me