8 months ago
Bengaluru, IndiaStaff+

Responsibilities

  • Lead critical production incident response, investigations, root-cause analysis, and resolution.
  • Architect and implement infrastructure improvements, automation frameworks, and observability systems.
  • Own initiatives that improve production resilience, reduce MTTR, and optimize infrastructure performance.
  • Mentor SRE team members, conduct code and design reviews, and develop engineering capabilities.
  • Lead blameless postmortems and improve incident response playbooks.
  • Collaborate with DevOps and engineering leadership on architecture and reliability standards.
  • Define and track SLIs, SLOs, and SLAs and implement monitoring strategies.
  • Participate in and help coordinate the global on-call rotation.

Requirements

  • 6+ years of hands-on AWS experience, including expert knowledge of EC2, ECS, EKS, RDS, S3, VPC, load balancing, CloudFormation, and multi-account strategies.
  • Strong leadership and mentorship experience leading technical initiatives and developing engineering talent.
  • Expertise in Linux systems administration and performance tuning.
  • Advanced production experience with Terraform and Ansible at scale.
  • Deep Kubernetes and ECS expertise, including cluster management, scaling, and troubleshooting.
  • Strong experience designing and implementing CI/CD pipelines using Jenkins, GitLab CI, or similar tools.
  • Advanced observability experience with CloudWatch, Prometheus, Grafana, ELK/EFK, Datadog, or equivalent platforms.
  • Expert networking skills covering DNS, load balancing, TLS/SSL, VPNs, service mesh architectures, and complex connectivity troubleshooting.
  • Proficiency in Python, Bash, or Go for automation and tool building.
  • Excellent communication and technical documentation skills.
  • Experience with DORA metrics, error budgets, toil reduction, and SRE practices.
  • Preferred: security and compliance experience with SOC2, ISO, or FedRAMP; open-source SRE/DevOps contributions; multi-region high-availability architectures; FinOps and cloud cost optimization; or GitOps practices with ArgoCD or Flux.

Benefits

  • Flexible workplace policies, employee resource groups, learning and development resources, career progression pathways, and community engagement initiatives.
  • Global employee wellness programs supporting physical, emotional, and financial well-being.
  • Benefits and perks vary by country.
  • The position is tagged as hybrid and is located in Bangalore.

Tech Stack

AnsibleAWSBashDatadogGitLab CI/CDGoGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Riverbed Technology, Inc.

About Riverbed Technology, Inc.

1,001-5,000 employees
Contact me