GrepJob
Okta

Staff Site Reliability Engineer

Okta
Apply
about 3 hours ago
Bengaluru, IndiaStaff+
H1B Sponsor

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and production services.
  • Participate in an on-call rotation supporting highly available customer-facing systems.
  • Lead incident response efforts and drive post-incident reviews focused on systemic improvements.
  • Define, measure, and improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Partner with engineering teams to improve service availability, scalability, performance, and resilience.
  • Continuously improve observability through metrics, logging, tracing, dashboards, and alerting.
  • Develop software, automation, and infrastructure using Go, Python, Terraform, and related technologies.
  • Eliminate operational toil through automation, tooling, and platform engineering.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Lead complex reliability initiatives spanning multiple engineering teams.
  • Guide engineers in adopting operational best practices and reliability engineering principles.
  • Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
  • Drive projects from conception through production rollout and long-term operational ownership.
  • Explore and apply AI-assisted engineering techniques to improve operational efficiency.

Requirements

  • Strong experience operating large-scale production services in AWS and/or GCP.
  • Deep expertise with Kubernetes in production environments.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
  • Strong understanding of cloud networking fundamentals including DNS, load balancing, ingress, and traffic management.
  • Experience with observability platforms, monitoring strategies, and production telemetry.
  • Understanding of cloud security fundamentals, IAM, and secure infrastructure design.
  • Demonstrated success leading complex engineering initiatives across multiple teams.

Benefits

  • Immersive, in-person onboarding experience designed to accelerate your impact.
  • Support for well-being and social impact initiatives.
  • Opportunities for talent development and fostering community connections.

Tech Stack

Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform