4 hours ago
Bengaluru, IndiaStaff+
Responsibilities
- Design, build, and operate large-scale cloud infrastructure and customer-facing production services.
- Participate in on-call rotations, lead incident response, and drive post-incident systemic improvements.
- Define and improve SLIs, SLOs, error budgets, availability, scalability, performance, resilience, and capacity planning.
- Develop software, automation, infrastructure, self-service platforms, and operational guardrails using Go, Python, Terraform, and related technologies.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Build observability through metrics, logging, tracing, dashboards, alerting, and production telemetry.
- Lead complex reliability initiatives across multiple engineering teams and drive projects from conception through production ownership.
- Mentor engineers, conduct design reviews and incident analysis, and influence technical direction through data-driven recommendations.
- Explore AI-assisted engineering and emerging technologies to improve troubleshooting, incident response, automation, and engineering productivity.
Requirements
- Strong experience operating large-scale production services in AWS and/or GCP.
- Deep production expertise with Kubernetes, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
- Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, traffic management, cloud security, IAM, secrets management, and secure infrastructure design.
- Experience with observability platforms, monitoring strategies, production telemetry, incident response, deployment strategies, and reliability engineering concepts.
- Demonstrated success leading cross-team engineering initiatives, mentoring engineers, and influencing technical direction.
- Experience working effectively in globally distributed engineering organizations across multiple time zones and cultures.
- Preferred experience includes operating SaaS platforms serving large-scale customer workloads, Kubernetes-based microservices, globally distributed production environments, GitOps, ArgoCD, and AI-assisted operational tooling.
Benefits
- The role is hybrid, with an immersive in-person onboarding experience.
- Okta highlights well-being support, social impact, talent development, and global community connection.
- The company has a global community spanning more than 20 offices worldwide.
Tech Stack
Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform
Categories
DevOpsSite Reliability
About Okta
Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.
