7 hours ago
Bengaluru, IndiaSenior
Responsibilities
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in an on-call rotation and lead incident response for highly available customer-facing systems.
- Define and improve SLIs, SLOs, error budgets, availability, scalability, performance, resilience, and capacity planning.
- Develop software, automation, infrastructure, self-service platforms, operational guardrails, and internal engineering tooling.
- Improve deployment safety and operational workflows through CI/CD and GitOps practices.
- Improve observability through metrics, logging, tracing, dashboards, alerting, and production telemetry.
- Modernize workloads and execute engineering projects from conception through production rollout and long-term ownership.
- Drive reliability initiatives, mentor engineers, conduct design reviews and incident analyses, and contribute to technical direction.
- Explore AI-assisted engineering and operational automation to reduce toil and improve incident response and engineering productivity.
Requirements
- Strong experience operating large-scale production services in AWS and/or GCP and deep expertise with Kubernetes in production.
- Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
- Extensive experience with Terraform and Helm or similar Infrastructure as Code technologies.
- Strong software engineering skills in Golang and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating and troubleshooting distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
- Strong understanding of cloud networking, DNS, load balancing, ingress, TLS, service networking, traffic management, observability, and monitoring strategies.
- Experience leading incident response and applying reliability engineering concepts including SLIs, SLOs, error budgets, and capacity planning.
- Strong understanding of CI/CD pipelines, deployment strategies, and automation-first operational practices.
- Understanding of cloud security fundamentals, IAM, secrets management, and secure infrastructure design.
- Demonstrated experience contributing to complex engineering initiatives, mentoring engineers, and collaborating across globally distributed organizations.
- Preferred qualifications include experience with large-scale SaaS platforms, Kubernetes-based microservices, globally distributed production environments, GitOps, ArgoCD, and AI-assisted operational tooling.
Benefits
- Hybrid role with an immersive in-person onboarding experience.
- Okta provides well-being support, social impact programs, talent development, and community-building opportunities.
- The company has a global community spanning more than 20 offices worldwide.
Tech Stack
Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformHelmKubernetesMySQLPostgreSQLPythonRedisSplunkTerraform
Categories
DevOpsSite Reliability
About Okta
Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.
