GrepJob
NICE

Specialist Cloud Site Reliability Engineer

NICE
Apply
about 3 hours ago
Pune, IndiaSenior / Staff+
H1B Sponsor

Responsibilities

  • Design and implement scalable, reliable systems across hybrid or multi-cloud environments.
  • Drive improvements in system uptime, latency, and service health metrics.
  • Build and manage infrastructure automation using Terraform, Helm, and Kubernetes.
  • Enhance CI/CD pipelines for safe and automated rollouts.
  • Own and enhance the observability stack with tools like Prometheus and Grafana.
  • Lead major incident response and root cause analysis.
  • Ensure platform-level security and compliance with standards.
  • Mentor junior SREs and contribute to technical roadmaps.

Requirements

  • 8+ years of experience with Kubernetes, EKS, ECS, and containerized workloads.
  • Expertise in AWS services such as EC2, Lambda, and RDS.
  • Proficiency with Terraform, Helm, Jenkins, and GitOps tools.
  • Deep understanding of observability frameworks and distributed monitoring.
  • Hands-on experience with Prometheus, Grafana, and OpenTelemetry.
  • Strong knowledge of Linux and system performance tuning.
  • Familiarity with Python, Go, or Shell scripting for automation.
  • Practical experience in incident response and on-call operations.

Benefits

  • Flexible work model with 2 days in the office and 3 days remote each week.
  • Opportunities for career growth across multiple roles and disciplines.
  • Collaborative and creative work environment.

Tech Stack

AWSGitHub ActionsGoGrafanaHelmJenkinsKubernetesPrometheusPythonTerraform

Categories