7 days ago
Base Salary
$165k - $227k/yr
Responsibilities
- Design, build, and operate large-scale cloud infrastructure and production services.
- Participate in global on-call rotations, incident response, and post-incident reviews.
- Define and improve SLIs, SLOs, error budgets, availability, scalability, performance, resilience, and capacity planning.
- Improve observability through metrics, logging, tracing, dashboards, alerting, and production telemetry.
- Develop software, automation, infrastructure, internal platforms, operational guardrails, and developer self-service tooling.
- Modernize workloads and improve deployment safety through Terraform, Helm, CI/CD, and GitOps practices.
- Ensure infrastructure and operational practices meet FedRAMP compliance and security requirements with continuous audit readiness.
- Collaborate on architecture and reliability initiatives, mentor engineers, conduct code reviews, and drive projects through production rollout and ongoing ownership.
- Explore AI-assisted engineering techniques for operational efficiency, troubleshooting, incident response, and automation.
Requirements
- Strong experience operating large-scale customer-facing production services in AWS and/or GCP.
- Deep production expertise with Linux and Kubernetes, including networking, storage, scheduling, scaling, and workload lifecycle troubleshooting.
- Extensive experience with infrastructure as code technologies such as Terraform and Helm.
- Strong software engineering skills in Go and/or Python.
- Experience building automation and internal engineering platforms.
- Experience operating distributed data platforms such as PostgreSQL, Redis, OpenSearch, MySQL, Cassandra, or similar technologies.
- Strong understanding of cloud networking fundamentals, including DNS, load balancing, ingress, TLS, service networking, and traffic management.
- Experience with observability platforms, monitoring strategies, production telemetry, incident response, SLIs, SLOs, error budgets, capacity planning, CI/CD pipelines, deployment strategies, and automation-first operations.
- Understanding of cloud security, IAM, secrets management, secure infrastructure design, and compliance standards such as FedRAMP, SOC2, and HIPAA.
- Demonstrated success contributing to complex engineering initiatives and influencing technical direction through expertise, partnership, and execution.
- Experience working with globally distributed engineering organizations and across multiple time zones and cultures.
- US Person status, defined as a US Citizen or Green Card Holder, is required for FedRAMP projects.
- Preferred qualifications include SaaS platform operations, Kubernetes-based microservices, globally distributed production environments, GitOps, ArgoCD, and AI-assisted operational tooling.
Benefits
- Hybrid work arrangement, indicated by the #LI-Hybrid designation.
- Immersive in-person onboarding designed to accelerate impact and build team connections.
- Health, dental, and vision insurance.
- 401(k) and flexible spending account.
- Paid leave, including PTO and parental leave.
- Equity and bonus opportunities may be available in addition to base salary.
- Well-being support, social impact programs, talent development, and community-building initiatives.
Tech Stack
Apache CassandraAWSDatadogGitGoGoogle Cloud PlatformGrafanaHelmKubernetesLinuxMySQLPostgreSQLPythonRedisRustSnowflakeSplunkTerraform
Categories
DevOpsSite Reliability
About Okta
Okta secures AI. Okta is The World’s Identity Company. Freeing everyone to safely use any technology—anywhere, on any device or app.
