Slate Magazine

Staff Software Engineer, Devops

Slate Magazine
Apply
1 day ago
Remote, United StatesStaff+

Base Salary

$156k - $234k/yr

Responsibilities

  • Architect and build Slate’s hybrid Kubernetes platform across Amazon EKS and self-managed clusters on virtualized infrastructure.
  • Own reference architectures for cluster provisioning, multi-cluster management, networking, identity, secrets, storage, and upgrades.
  • Evolve the multi-account AWS environment, including governance, guardrails, IAM, networking, DNS, and hybrid connectivity.
  • Establish on-call rotation, SLOs, error budgets, alerting standards, incident response, postmortems, and observability practices.
  • Provide teams with CI/CD templates, deployment patterns, reusable infrastructure libraries, and self-service environments.
  • Support workloads across digital services, data pipelines, manufacturing systems, plant-floor environments, and AI or agentic systems.
  • Build in least-privilege access, policy guardrails, secrets management, software supply chain controls, and cloud cost discipline.
  • Write production infrastructure and code, lead design reviews, mentor engineers, and maintain documentation and runbooks.

Requirements

  • 10+ years of experience in DevOps, site reliability, platform, or infrastructure engineering, including several years setting technical direction for production platforms.
  • Extensive hands-on experience designing and operating production AWS environments across multiple accounts, including governance, IAM, VPC networking, hybrid connectivity, DNS, EKS, containers, serverless, and managed data services.
  • Proven experience running Kubernetes in production on EKS and self-managed clusters on virtual machines or bare metal, including lifecycle management, CNI networking, ingress, storage, RBAC, and upgrades.
  • Strong infrastructure-as-code experience and strong Python skills; AWS CDK is strongly preferred, and the ability to make targeted changes in TypeScript is required.
  • Experience building delivery pipelines with GitHub Actions or similar tools and deploying to Kubernetes with GitOps tooling, Helm, and Kustomize.
  • Track record establishing SLOs, observability, incident management, on-call practices, and reliability engineering processes.
  • Solid understanding of cloud and on-premises networking, including routing, VPNs, DNS, load balancing, certificates, and firewalls.
  • Ability to drive architectural decisions across teams without direct authority, write clear design documents, and mentor engineers.
  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
  • AWS Professional or Specialty certification, Rancher or comparable multi-cluster management experience, manufacturing or OT exposure, GPU or machine learning workload experience, VMware familiarity, policy-as-code experience, service mesh experience, or GitHub Enterprise administration are preferred.

Benefits

  • Medical, dental, vision, life insurance, disability insurance, vacation, and 401(k) benefits are offered.
  • The role may include participation in an equity program and/or discretionary annual incentive program, subject to applicable rules.
  • Applicants must be authorized to work permanently in the United States; visa sponsorship is unavailable.

Tech Stack

Argo CDAWSGitHub ActionsGrafanaHelmKubernetesPrometheusPythonRancherTypeScript

Categories

DevOpsSite Reliability
Slate Magazine

About Slate Magazine

11-50 employees

Slate is a digital magazine that covers politics, culture, technology, and current events through reported articles, analysis, newsletters, and a large slate of podcasts. Founded in 1996, it is published by The Slate Group, a unit of Graham Holdings, with offices in New York and Washington, D.C. Its business model combines advertising with reader revenue, including the Slate Plus membership program for bonus content and ad-free podcasts.

Contact me