22 hours ago
Remote, WorldwideSenior
Base Salary
$85k - $155k/yr
Responsibilities
- Own global Kubernetes clusters across regions and cloud providers, including provisioning, upgrades, multi-tenancy, resource and network policy, and failure isolation.
- Own the internal cluster supporting CI/CD runners and build platforms, including capacity, isolation, scaling, and build performance.
- Manage shared infrastructure such as databases, object storage, container registries, and build platforms, including provisioning, access models, and cost.
- Own the lifecycle of shared platform services, including internal PKI, vaults, and Kubernetes operators.
- Design and exercise disaster-recovery capabilities, including backup and restore, replication, failover, RTO/RPO targets, and recovery testing.
- Design and publish reusable Terraform/OpenTofu modules with built-in correctness and policy controls.
- Own the observability stack and define instrumentation standards for other engineering teams.
- Build production tooling in Go or Python.
- Mentor engineers through code reviews, design reviews, and technical guidance.
- Reduce infrastructure cost and drift, and participate in the Platform Infrastructure on-call rotation.
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field.
- At least 5 years of professional experience in infrastructure, platform, DevOps, or SRE engineering with production ownership.
- Deep production experience operating Kubernetes clusters.
- Strong Terraform or OpenTofu experience at scale, including module design, state management, and policy-as-code.
- Strong programming ability in Go, Python, or a comparable language.
- Experience operating stateful shared services or infrastructure in production, such as message brokers, databases, or certificate authorities.
- Deep Linux fundamentals and system-level debugging under load.
- Understanding of TCP/IP, DNS, HTTP, load balancing, and service-to-service communication.
- Hands-on experience with at least one major cloud provider and ability to work across others.
- Sound judgment regarding blast radius and reversible changes.
- Preferred experience on internal or developer-platform teams serving other engineers.
- Preferred experience with Scalr or comparable IaC platforms, OPA, ArgoCD, Flux, Helm, Kustomize, or operator development.
- Preferred experience managing self-hosted CI runners and working with Prometheus/Thanos, exporters, instrumentation, telemetry, logging, distributed tracing, and monitoring architectures.
- Preferred experience building highly available, multi-tenant SaaS platforms and influencing engineering strategy while mentoring technical talent.
Benefits
- The position includes participation in production on-call.
- The interview process includes a coding challenge, hiring-manager screening, and technical interviews.
- The role is based in a fast-moving, collaborative organization with support and shared accountability.
- A background screening may include criminal-record, motor-vehicle, or drug screening checks depending on position requirements.
About Alkira
Alkira builds a Network‑infrastructure‑as‑a‑Service platform for enterprises to design, deploy, and operate connectivity across multicloud, data centers, branches, and partner environments. Delivered as a cloud service, it centralizes routing, segmentation, and security controls and integrates with AWS, Azure, GCP, and third‑party firewalls. Founded in 2018 and headquartered in San Jose, Alkira is a Lumen Technologies company.
