2 days ago
Dublin, IrelandStaff+
Responsibilities
- Own the architecture and evolution of the Dublin/EMEA Kubernetes platform, including cluster strategy, multi-tenancy, networking, and security posture.
- Design and roll out self-service platform capabilities, golden paths, and self-healing patterns for internal engineering teams.
- Drive reliable, scalable infrastructure for AI-driven internal workflows, including model serving, GPU scheduling, and orchestration.
- Resolve complex production incidents, capacity and scaling problems, and cross-system failure modes.
- Participate in a follow-the-sun on-call rotation with SRE teams in the US and India and lead root-cause analysis and postmortems.
- Set technical standards for SLIs, SLOs, error budgets, and on-call practices.
- Improve infrastructure-as-code SDLC processes, CI/CD maturity, and change and release management.
- Mentor engineers through design reviews, architecture discussions, and hands-on pairing.
- Partner with architects, security, and product engineering teams on infrastructure, reliability, security, and delivery decisions.
- Represent the Dublin/EMEA team in global platform architecture discussions and influence technical direction across distributed teams.
Requirements
- 8+ years of experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering with staff-level technical leadership.
- Deep hands-on expertise in production-scale Kubernetes, including architecture, multi-tenancy, networking, security, and day-2 operations.
- Experience operating large-scale internal infrastructure platforms in a public cloud, preferably AWS.
- Strong expertise in cloud-native architecture, Terraform, and CI/CD pipelines.
- Experience operating infrastructure for AI/ML or agentic workflows, including model serving, GPU scheduling, or orchestration frameworks.
- Deep experience with observability and monitoring tools such as Grafana, Splunk, or APM platforms.
- Ability to influence technical direction across teams and geographies without formal authority.
- Strong verbal and written communication, mentoring, and technical-alignment skills.
- Computer Science degree or related field, or equivalent experience.
- Experience building internal developer platforms with a platform-as-a-product mindset is preferred.
- Experience with AI agent orchestration, vector databases, model routing, or inference optimization infrastructure is preferred.
- Multi-cloud experience with AWS plus Azure or GCP is preferred.
- Production experience with Istio, Linkerd, ArgoCD, or Flux is preferred.
- Kubernetes certifications such as CKA, CKS, or CKAD, or equivalent cloud certifications, are preferred.
- Experience administering enterprise-scale SCM platforms such as GitHub or GitLab is preferred.
Benefits
- Hybrid work arrangement in Dublin/EMEA, with an immersive in-person onboarding experience.
- Equity where applicable, comprehensive healthcare coverage, financial benefits, paid time off, and parental leave.
- Opportunities for talent development, social impact, and connection across Okta’s global community.
Tech Stack
Categories
DevOpsSite Reliability
About Okta
Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.
