13 hours ago
Remote, United States or Remote, CanadaSenior
Base Salary
$140k - $175k/yr
Responsibilities
- Rearchitect and migrate AWS and Azure infrastructure to OpenTofu/Terraform and Ansible, including reusable modules, remote state, GitOps pipelines, policy-as-code, and cloud governance.
- Build and operate observability systems covering metrics, logs, traces, dashboards, OpenTelemetry instrumentation, SLIs, SLOs, error budgets, and alerting.
- Establish SRE practices including incident runbooks, post-incident reviews, chaos engineering, capacity planning, and operational dashboards.
- Deploy and operate an Internal Developer Platform and developer portal with service catalogs, scaffolding templates, runbooks, API documentation, and self-service provisioning.
- Own CI/CD platform capabilities, reusable workflow libraries, container build and scanning pipelines, environment promotion, and security scanning.
- Operate EKS and/or AKS clusters and establish standards for Helm, admission controllers, RBAC, network policies, autoscaling, and service mesh.
- Support workload migrations, embed security controls, write technical documentation and ADRs, participate in architecture reviews, and mentor junior platform engineers.
Requirements
- 5+ years of platform, infrastructure, or DevOps engineering experience with direct production ownership on AWS and/or Azure.
- Deep OpenTofu/Terraform experience covering module authoring, state management, workspace strategy, remote backends, and CI/CD integration.
- Strong Kubernetes operations experience with EKS and/or AKS, Helm, admission controllers, RBAC, network policies, and autoscaling.
- Hands-on experience with at least two observability technologies such as Prometheus, Grafana, Loki, Tempo, Datadog, or OpenTelemetry, including SLI/SLO definition and alert engineering.
- Experience authoring GitHub Actions pipelines, designing reusable workflows, and owning container build and scanning pipelines.
- Experience with GitOps tools such as ArgoCD or Flux and with Internal Developer Platforms, Backstage, developer portals, service catalogs, scaffolding, or self-service provisioning.
- Strong security-first practices involving policy-as-code, IaC scanning, secrets management, container hardening, and shift-left security.
- Strong communication and documentation skills, with the ability to present architecture decisions to engineering peers and leadership.
- Preferred qualifications include Terramate, progressive delivery, AWS FIS or Chaos Monkey, error budget management, incident command, Istio or Linkerd, Kubecost, CloudHealth, AI/ML infrastructure, Python or Go, and relevant AWS, Azure, Kubernetes, or Terraform certifications.
Benefits
- Competitive compensation with a stated base pay range of $140,000-$175,000.
- Health, dental, and vision insurance.
- 401(k) plan with company match.
- Generous paid time off.
- Flexible remote, hybrid, or in-office work arrangements.
- Inclusive and collaborative workplace culture.
Tech Stack
Categories
DevOpsSite Reliability
