Okta

Staff Site Reliability Engineer, Federal (TS/SCI)

Okta
Apply
2 months ago
Washington, DC, USAStaff+
H1B sponsor

Base Salary

$174k - $238k/yr

Responsibilities

  • Design, build, and operate large-scale cloud infrastructure and highly available customer-facing production services.
  • Participate in on-call rotations, lead incident response, and drive post-incident reviews and systemic improvements.
  • Define and improve SLIs, SLOs, error budgets, availability, scalability, performance, resilience, and capacity planning.
  • Develop software, automation, infrastructure, self-service platforms, operational guardrails, and internal tooling.
  • Improve deployment safety and operational workflows through CI/CD and GitOps practices.
  • Improve observability through metrics, logging, tracing, dashboards, alerting, and production telemetry.
  • Lead reliability initiatives across multiple engineering teams and guide adoption of operational best practices.
  • Mentor engineers through technical collaboration, design reviews, incident analysis, and knowledge sharing.
  • Influence architecture and operational decisions and own projects from conception through production rollout and ongoing operations.
  • Explore AI-assisted engineering and emerging technologies to reduce operational toil and improve incident response and productivity.

Requirements

  • Active U.S. TS/SCI clearance with Full Scope Poly.
  • Proven experience navigating Federal and DoD compliance frameworks, specifically FedRAMP and Impact Level 6 (IL6).
  • Strong experience operating large-scale production services in AWS and/or GCP and deep production Kubernetes expertise.
  • Experience troubleshooting Kubernetes networking, storage, scheduling, scaling, and workload lifecycle issues.
  • Extensive experience with Infrastructure as Code technologies such as Terraform and Helm.
  • Strong software engineering skills in Golang and/or Python.
  • Experience building automation and internal engineering platforms.
  • Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, or Cassandra.
  • Strong understanding of cloud networking, observability, monitoring, production telemetry, CI/CD pipelines, deployment strategies, and reliability engineering concepts.
  • Experience leading incident response, complex engineering initiatives, and operational improvements across multiple teams.
  • Experience mentoring engineers and influencing technical direction in globally distributed organizations.
  • Preferred qualifications include operating large-scale SaaS platforms, Kubernetes-based microservices, globally distributed production environments, GitOps and ArgoCD, and AI-assisted operational tooling.
  • Must be on U.S. soil and able to obtain and maintain a U.S. security clearance as required by U.S. Government contracts.

Benefits

  • Health, dental, and vision insurance.
  • 401(k), flexible spending account, and paid leave including PTO and parental leave.
  • Immersive in-person onboarding experience.
  • Employees must be located on U.S. soil and maintain the required U.S. security clearance.

Categories

DevOpsSite Reliability
Okta

About Okta

5,001-10,000 employees

Okta builds cloud-based identity and access management for enterprises and developers, including single sign-on, multi-factor authentication, and lifecycle management. It sells subscription SaaS as the Okta Workforce Identity and Customer Identity Clouds; the latter incorporates Auth0, acquired in 2021. Founded in 2009 and headquartered in San Francisco, Okta is a public company traded on Nasdaq.

Contact me