10 days ago
Responsibilities
- Own and improve incident response processes in collaboration with other teams.
- Participate independently in rotational on-call coverage for critical systems 24/7.
- Use monitoring tools to identify, resolve, and escalate production issues.
- Implement infrastructure resilience, monitoring, and alerting improvements.
- Set and achieve long-term performance, reliability, and scalability goals for Auth0 systems.
Requirements
- At least 3 years of experience as a Site Reliability Engineer or in a Cloud Operations/DevOps role.
- At least 2 years of experience using Go, shell scripting, and Terraform.
- At least 2 years of experience as a software developer in a SaaS environment.
- At least 3 years of experience supporting large-scale, mission-critical applications in production.
- Knowledge of AWS, Azure, databases, containers, web technologies, networking, SLIs, SLOs, error budgets, and SLAs.
- Knowledge of Datadog or another observability platform is preferred.
- Strong technical communication, systematic problem-solving, ownership, automation, teamwork, and remote-work capabilities.
Benefits
- Remote work environment with self-directed responsibilities.
- Immersive in-person onboarding designed to accelerate impact and build team connections.
- Programs supporting employee well-being, social impact, talent development, and community connection.
About Okta
Okta secures AI. Okta is The World’s Identity Company. Freeing everyone to safely use any technology—anywhere, on any device or app.
