Okta

Staff Site Reliability Engineer

Okta
Apply
18 days ago
Bengaluru, IndiaStaff+

Responsibilities

  • Design, build, operate, and improve large-scale cloud infrastructure and highly available customer-facing production services.
  • Participate in on-call rotations, lead incident response, and drive post-incident reviews and systemic reliability improvements.
  • Define and improve SLIs, SLOs, error budgets, availability, scalability, performance, resilience, and capacity planning.
  • Develop software, automation, infrastructure, self-service platforms, operational guardrails, and internal engineering tooling using Golang, Python, Terraform, and related technologies.
  • Improve observability through metrics, logging, tracing, dashboards, alerting, and production telemetry.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Lead complex reliability initiatives across multiple engineering teams and guide adoption of operational best practices.
  • Mentor engineers, conduct design reviews and incident analysis, and influence architecture and operational decisions.
  • Explore AI-assisted engineering and emerging technologies to improve incident response, troubleshooting, automation, and engineering productivity.

Requirements

  • Strong experience operating large-scale production services in AWS and/or GCP.
  • Deep production expertise with Kubernetes, including troubleshooting networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Extensive experience with infrastructure as code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
  • Strong understanding of cloud networking fundamentals, observability platforms, monitoring strategies, reliability engineering, CI/CD, deployment strategies, and automation-first operations.
  • Experience leading incident response, operational improvements, and complex engineering initiatives across multiple teams.
  • Understanding of cloud security fundamentals, IAM, secrets management, secure infrastructure design, and operational controls in regulated or security-sensitive environments.
  • Experience mentoring engineers and working effectively within globally distributed engineering organizations.
  • Preferred experience includes operating SaaS platforms at scale, Kubernetes-based microservices, globally distributed production environments, GitOps, ArgoCD, and AI-assisted operational tooling.

Benefits

  • Hybrid work arrangement with an immersive, in-person onboarding experience.
  • Well-being support, social impact opportunities, talent development, and community connection.
  • Access to a global community spanning more than 20 offices worldwide.

Tech Stack

Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform

Categories

DevOpsSite Reliability
Okta

About Okta

5,001-10,000 employees

Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.

Contact me