Empower

Sr Engineer Site Reliability

Empower
Apply
10 hours ago
Remote, India or Bengaluru, IndiaMid Level / Senior

Responsibilities

  • Design and implement highly available and fault-tolerant systems supporting critical financial transactions.
  • Architect AWS infrastructure solutions optimized for cost, performance, and reliability.
  • Lead complex incident response efforts and coordinate service restoration across teams.
  • Drive postmortems for high-severity incidents and ensure action items are completed.
  • Establish and track SLOs and SLIs for key services.
  • Design disaster recovery strategies and business continuity plans.
  • Build advanced Terraform infrastructure-as-code solutions using modules, workspaces, and state management.
  • Architect and optimize multi-cluster EKS environments, including pod and cluster autoscaling.
  • Design observability strategies, dashboards, metrics, and alerting using Datadog and Splunk.
  • Implement canary and blue-green deployments within GitOps workflows.
  • Build automation that reduces operational toil and improves team efficiency.
  • Partner with development teams on reliability improvements, design reviews, and architectural guidance.
  • Mentor junior and intermediate SREs, conduct code reviews, and provide technical coaching.
  • Evangelize SRE practices and contribute to reliability and scalability decisions.
  • Participate in on-call rotations and reduce the associated operational burden.
  • Implement zero-trust infrastructure controls and support financial-services compliance requirements.
  • Conduct infrastructure security reviews, audit preparation, and responses to compliance inquiries.

Requirements

  • Bachelor's degree in Computer Science, Information Systems, or a similar field, or equivalent experience.
  • 4–7 years of Site Reliability Engineering or equivalent experience operating large-scale production systems.
  • Deep AWS expertise and hands-on experience with broad AWS services and architectural patterns.
  • Advanced Kubernetes knowledge, including custom resources, operators, and cluster federation concepts.
  • Expert Terraform proficiency, including module development, state management, and complex workflow orchestration.
  • Strong Python and/or Go programming skills for production-quality tools and services.
  • Production-scale observability experience with Datadog, Splunk, or similar platforms.
  • Experience establishing and maintaining enterprise-scale CI/CD pipelines.
  • Strong understanding of GitOps and experience with ArgoCD or Flux.
  • Experience leading complex incident response and conducting thorough postmortems.
  • Strong understanding of networking, security, and infrastructure design patterns.
  • Experience mentoring engineers and conducting technical training.
  • Preferred experience in financial services or payments, compliance frameworks such as SOC 2, PCI DSS, or FINRA, AWS certifications, CKA or CKAD certifications, service mesh implementations, chaos engineering, FinOps, and open-source SRE or DevOps contributions.

Benefits

  • Flexible work environment, fluid career paths, and opportunities for internal mobility.
  • Emphasis on purpose, well-being, work-life balance, inclusion, and community volunteering.
  • Participation in on-call rotations is expected.

Tech Stack

AWSConsulDatadogGitLab CI/CDGoGrafanaHelmIstioJenkinsKubernetesPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
Empower

About Empower

10,000+ employees
Contact me