Okta

Staff Site Reliability Engineer

Okta
Apply
2 days ago
Dublin, IrelandStaff+

Responsibilities

  • Own the architecture and evolution of the Dublin/EMEA Kubernetes platform, including cluster strategy, multi-tenancy, networking, and security posture.
  • Design and roll out self-service platform capabilities, golden paths, and self-healing patterns for internal engineering teams.
  • Drive reliable, scalable infrastructure for AI-driven internal workflows, including model serving, GPU scheduling, and orchestration.
  • Resolve complex production incidents, capacity and scaling problems, and cross-system failure modes.
  • Participate in a follow-the-sun on-call rotation with SRE teams in the US and India and lead root-cause analysis and postmortems.
  • Set technical standards for SLIs, SLOs, error budgets, and on-call practices.
  • Improve infrastructure-as-code SDLC processes, CI/CD maturity, and change and release management.
  • Mentor engineers through design reviews, architecture discussions, and hands-on pairing.
  • Partner with architects, security, and product engineering teams on infrastructure, reliability, security, and delivery decisions.
  • Represent the Dublin/EMEA team in global platform architecture discussions and influence technical direction across distributed teams.

Requirements

  • 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering with staff-level technical leadership.
  • Deep hands-on expertise in production-scale Kubernetes, including architecture, multi-tenancy, networking, security, and day-2 operations.
  • Experience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS.
  • Strong expertise in cloud-native architecture, Terraform, and CI/CD pipelines.
  • Experience operating infrastructure for AI/ML or agentic workflows, including model serving, GPU scheduling, or orchestration frameworks.
  • Deep experience with observability and monitoring tools such as Grafana, Splunk, or APM platforms.
  • Ability to influence technical direction across teams and geographies without formal authority.
  • Strong verbal and written communication, mentoring, and technical-alignment skills.
  • Computer Science degree or related field, or equivalent experience.
  • Experience building internal developer platforms with a platform-as-a-product mindset is preferred.
  • Experience with AI agent orchestration, vector databases, model routing, or inference optimization infrastructure is preferred.
  • Multi-cloud experience with AWS plus Azure or GCP is preferred.
  • Production experience with Istio, Linkerd, ArgoCD, or Flux is preferred.
  • Kubernetes certifications such as CKA, CKS, or CKAD, or equivalent cloud certifications, are preferred.
  • Experience administering enterprise-scale SCM platforms such as GitHub or GitLab is preferred.

Benefits

  • Hybrid work arrangement in Dublin/EMEA, with an immersive in-person onboarding experience.
  • Equity where applicable, comprehensive healthcare coverage, financial benefits, paid time off, and parental leave.
  • Opportunities for talent development, social impact, and connection across Okta’s global community.

Categories

DevOpsSite Reliability
Okta

About Okta

5,001-10,000 employees

Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.

Contact me