Grubhub

Staff Site Reliability Engineer

Grubhub
Apply
1 day ago

Base Salary

$209k - $217k/yr

Responsibilities

  • Architect resilient, self-healing systems and co-own critical production service design.
  • Own multi-region resilience, including active-standby architecture, regional failover, runbooks, drills, RTO/RPO targets, and data-layer replication.
  • Design, roll out, and maintain AWS infrastructure as code using Terraform, Terraspace, Helm, and Helmfile.
  • Own the Kubernetes platform, including EKS lifecycle, controller and add-on upgrades, ingress, gateways, and autoscaling.
  • Operate the observability platform across logging, metrics, tracing, alerting, signal quality, and cost.
  • Drive reliability improvements using SLOs and telemetry, and develop seasonal scaling and capacity strategies.
  • Own cloud cost accountability through right-sizing, reservations, and savings tracking.
  • Build and maintain CI/CD pipelines and deployment tooling.
  • Operate segregated clusters and access paths in PCI-scoped environments.
  • Lead incident response, postmortems, failure analysis, service reviews, and architecture reviews.
  • Set technical direction and build consensus across engineering teams in multiple time zones.

Requirements

  • Demonstrated Staff-level scope owning a production platform end to end and driving technical direction across teams without formal authority.
  • Typically 8+ years of experience in SRE, DevOps, or infrastructure engineering.
  • Deep Infrastructure as Code experience with Terraform, including modules, state, and multi-environment rollouts; Terraspace or a similar wrapper is expected.
  • Operator-level Kubernetes experience with EKS, Helm, Helmfile, controllers, add-ons, ingress or gateways, and HPA/KEDA autoscaling.
  • Experience designing multi-region AWS architectures, failover systems, and dependent data-layer replication.
  • Deep knowledge of CI/CD tools such as Jenkins or GitHub Actions.
  • Software engineering experience in Python, Go, or a similar object-oriented language.
  • Proficiency with MySQL, MongoDB/Atlas, Redis/ElastiCache, and message brokers such as RabbitMQ/AmazonMQ and SQS.
  • Experience with distributed monitoring, SLOs, metrics, tracing, and log pipelines.
  • Strong knowledge of AWS compute and containers, storage, Linux, and networking.
  • Experience with highly trafficked web-based services and compliance-scoped environments such as PCI.
  • Strong technical writing, documentation, communication, consensus-building, incident response, and architecture review skills.

Benefits

  • Equity and a 401(k) are included in the compensation package.
  • Medical, dental, and vision plan options are available.
  • Company-paid short- and long-term disability coverage is provided.
  • Paid time off includes flexible time off for exempt employees, paid vacation for non-exempt employees, and paid sick leave as required by law.
  • Paid parental leave, discounted meals, and exclusive perks across the Wonder family of brands are offered.
  • Benefit eligibility, effective dates, and plan options vary by employment classification and location.
  • The role uses geographic-specific salary structures, with the stated base salary for New York.

Tech Stack

AWSGitHub ActionsGoHelmJenkinsKubernetesLinuxMongoDBMySQLPythonRabbitMQRedisTerraform

Categories

DevOpsSite Reliability
Grubhub

About Grubhub

5,001-10,000 employees

Grubhub builds an online marketplace and logistics network for ordering takeout and delivery from local restaurants across the U.S. It earns commissions and delivery fees, and offers consumer subscriptions and corporate solutions. Founded in 2004 and headquartered in Chicago, it is part of Wonder Group and lists hundreds of thousands of merchants across 4,000+ cities.

Contact me