10 days ago
Responsibilities
- Define reliability strategy, architecture, standards, objectives, and operational guardrails for critical services and platforms.
- Own reliability architecture and operational excellence for the Spera/ISPM product area and contribute to broader EPG reliability initiatives.
- Lead scalability, resilience, performance, architectural modernization, and platform maturity programs.
- Design, build, and operate large-scale cloud infrastructure and production services.
- Develop automation, tooling, internal platforms, and self-service operational capabilities using Go, Python, Terraform, and related technologies.
- Improve deployment safety, operational workflows, platform consistency, and developer productivity through GitOps and Infrastructure-as-Code practices.
- Establish observability, monitoring, incident management, operational readiness, and reliability engineering best practices.
- Support highly available customer-facing production systems through an on-call rotation and lead major incident response efforts.
- Mentor Staff and Senior engineers, lead technical and design reviews, and build consensus across distributed teams.
- Explore and champion AI-assisted and agentic systems for troubleshooting, root-cause analysis, incident response, and operational automation.
Requirements
- Extensive experience designing and operating large-scale production systems in AWS and/or GCP.
- Deep production expertise with Kubernetes and Kubernetes-based reliability strategies, including networking, storage, scheduling, scaling, and workload lifecycle challenges.
- Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python, with experience building internal platforms, developer tooling, and operational automation.
- Deep understanding of distributed systems, cloud-native application design, cloud networking, observability, and multi-region architectures.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
- Strong knowledge of reliability principles including SLIs, SLOs, error budgets, capacity planning, resilience engineering, CI/CD, GitOps, and deployment strategies.
- Demonstrated success leading complex technical initiatives, reliability transformations, and architectural modernization across multiple teams and organizations.
- Experience with customer-facing production systems at scale, major incident response, and long-term operational improvement.
- Strong understanding of cloud security fundamentals, IAM, secrets management, secure infrastructure design, and operational controls.
- Ability to influence technical direction without direct authority, work across globally distributed organizations, mentor technical leaders, and translate business objectives into measurable reliability outcomes.
- Experience or strong interest in AI-assisted engineering and operational automation.
- Preferred qualifications include operating SaaS platforms serving millions of users, supporting globally distributed environments, leading platform or reliability transformations, and implementing AI-assisted operational tooling or agentic workflows.
Benefits
- Hybrid work arrangement with an immersive, in-person onboarding experience.
- Opportunities for professional development, talent development, and connection across Okta’s global community.
- Well-being support, social impact programs, and community-building initiatives.
- Global workplace spanning more than 20 offices worldwide.
Tech Stack
Apache CassandraArgo CDAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform
Categories
Site Reliability
About Okta
Okta secures AI. Okta is The World’s Identity Company. Freeing everyone to safely use any technology—anywhere, on any device or app.
