ISS Governance

Senior Site Reliability Engineer

ISS Governance
Apply
1 month ago
Mumbai, IndiaSenior

Responsibilities

  • Design, expand, and optimize Kubernetes-based deployment architectures for microservices.
  • Enhance monitoring and observability using Grafana, Prometheus, and related tools.
  • Ensure consistent, actionable logging integrated with log-indexing infrastructure.
  • Write and maintain code for build automation, tooling, and operational workflows.
  • Build and improve Docker images and manage workloads across Kubernetes clusters.
  • Partner with DevOps and engineering teams on deployment practices, operational processes, and automation.
  • Develop documentation, runbooks, and escalation procedures while championing reliability best practices.

Requirements

  • Bachelor’s degree in computer science or a related technical field, or equivalent practical experience.
  • At least 4 to 5+ years of programming experience in one or more of Java, Python, Go, or Django.
  • Deep experience troubleshooting large-scale distributed systems.
  • Strong interest in software reliability engineering, automation, and resilient systems.
  • Strong experience with Linux/Unix environments and system internals.
  • Clear communication and effective cross-functional collaboration skills.

Benefits

  • The role is based in Mumbai, Goregaon East, India.
  • Shift hours are 1 PM IST to 10 PM IST.
  • ISS STOXX offers a culture focused on diversity, collaboration, professional growth, and personal development.

Tech Stack

DjangoDockerGoGrafanaJavaKubernetesLinuxPrometheusPython

Categories

Site Reliability
Contact me