GrepJob
Okta

Staff Site Reliability Engineer - Ecosystem

Okta
Apply
about 3 hours ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Design, build, and operate scalable and secure infrastructure across AWS and GCP.
  • Lead reliability and modernization initiatives, including container platform migrations.
  • Serve as a technical authority in Kubernetes and cloud infrastructure.
  • Partner with development teams to architect microservice-based applications.
  • Implement and manage infrastructure as code using Terraform and Ansible.
  • Drive improvements in observability, performance, and cost efficiency.
  • Champion SRE best practices and conduct blameless postmortems.
  • Lead complex technical projects from conception to completion.
  • Mentor engineers and foster a culture of reliability and automation.
  • Collaborate with security and compliance partners to ensure best practices.

Requirements

  • 8+ years in SRE, DevOps, or Infrastructure Engineering roles.
  • 3–5 years of experience with Kubernetes (EKS/GKE) in production.
  • 3–5 years of experience with AWS and GCP.
  • 3–5 years using Terraform for multi-cloud infrastructure management.
  • 3+ years of coding experience in Python, Go, or similar languages.
  • Proven track record leading migration projects and enabling microservice architectures.
  • Experience implementing SLOs/SLIs and improving operational resilience.
  • Strong Linux and security fundamentals.
  • Bachelor’s degree in Computer Science or equivalent experience.

Benefits

  • Immersive, in-person onboarding experience.
  • Support for employee well-being.
  • Opportunities for social impact and community connection.

Tech Stack

AnsibleAWSGitLab CI/CDGoGoogle Cloud PlatformGrafanaKubernetesLinuxMySQLPostgreSQLPrometheusPythonRedisSpinnakerTerraform

Categories