1 day ago
Hyderābād, IndiaMid Level
H1B Sponsor

Responsibilities

  • Deploy, manage, and scale distributed platforms across multiple geographic regions.
  • Design and maintain Kubernetes-based infrastructure for large-scale applications.
  • Build and manage Helm charts for repeatable application deployments.
  • Monitor system health with Grafana dashboards and metrics and proactively resolve issues.
  • Improve reliability, performance, scalability, and infrastructure through automation and best practices.
  • Collaborate with development teams on CI/CD and production readiness.
  • Implement observability, alerting, and incident response processes.
  • Troubleshoot production issues, perform root cause analysis, and maintain incident-response runbooks.

Requirements

  • Require 4–5 years of experience in Site Reliability Engineering, DevOps, or similar roles.
  • Require strong hands-on Kubernetes experience in production environments.
  • Require strong infrastructure-as-code experience with Terraform and Git.
  • Require strong AWS experience, including EKS, VPC, S3, ECR, and IAM roles.
  • Require solid experience creating Helm charts for application deployment.
  • Require strong Bash scripting and tooling experience.
  • Require experience with large-scale distributed systems and high-availability architectures.
  • Require understanding of containerization, microservices, and cloud-native ecosystems.
  • Require experience with CI/CD pipelines and automation tools.
  • Require production debugging and problem-solving skills.
  • Prefer proficiency in Golang or Python.
  • Prefer experience building and managing Grafana dashboards and metrics.
  • Prefer knowledge of monitoring and observability stacks such as Prometheus and Loki.
  • Prefer experience with multi-region deployments and global infrastructure.

Benefits

  • Competitive compensation and comprehensive benefits.
  • Flexible work environment.
  • Annual wellness and community outreach days.
  • Recognition for contributions.
  • Global collaboration and networking opportunities.

Tech Stack

AWSBashGitGoGrafanaHelmKubernetesPrometheusPythonTerraform

Categories

Site Reliability
Proofpoint

About Proofpoint

1,001-5,000 employees

Intent-based protection for every human and every AI agent, across all data.

Contact me