1 day ago
Tel Aviv-Yafo, IsraelStaff+
Responsibilities
- Architect resilient, self-healing systems and co-own critical production service design.
- Own multi-region resilience, including active-standby architecture, regional failover, runbooks, drills, RTO/RPO targets, and data-layer replication.
- Own AWS infrastructure as code using Terraform, Terraspace, Helm, and Helmfile across multiple environments.
- Own the Kubernetes platform, including EKS lifecycle, controller and add-on upgrades, ingress, gateways, and autoscaling.
- Own observability logging, metrics, tracing, and alerting pipelines while improving signal quality and controlling cost.
- Drive reliability, scaling, capacity planning, cloud cost optimization, and right-sizing for seasonal traffic.
- Build and maintain CI/CD pipelines and deployment tooling.
- Operate within PCI-scoped environments with segregated clusters and access paths.
- Lead incident response, postmortems, failure analysis, service reviews, and architecture reviews.
- Set technical direction and build consensus across engineering teams in multiple time zones.
Requirements
- Typically 8+ years of experience in SRE, DevOps, or infrastructure engineering, with demonstrated Staff-level scope and end-to-end production platform ownership.
- Deep experience with Terraform, including modules, state, and rollouts across multiple environments, with Terraspace or a similar wrapper.
- Operator-level Kubernetes experience with EKS, Helm, Helmfile, controllers, add-ons, ingress or gateways, and HPA/KEDA autoscaling.
- Experience designing multi-region AWS architectures, failover systems, and replicated data layers.
- Experience with CI/CD tools such as Jenkins or GitHub Actions.
- Software engineering experience in Python, Go, or a similar object-oriented language.
- Proficiency with MySQL, MongoDB/Atlas, Redis/ElastiCache, and message brokers such as RabbitMQ/AmazonMQ and SQS.
- Experience with microservice architecture, application design, distributed monitoring, SLOs, metrics, tracing, and log pipelines.
- Strong knowledge of AWS compute and containers, storage, Linux, networking, and cloud fundamentals.
- Experience operating in PCI or equivalent compliance-scoped environments and with highly trafficked web services.
- Strong technical writing, documentation, communication, technical direction, and consensus-building skills.
Benefits
- Competitive compensation package with equity and a 401(k).
- Medical, dental, and vision plan options.
- Company-paid short- and long-term disability coverage.
- Paid time off, including flexible time off for exempt employees, paid vacation for non-exempt employees, and paid sick leave as required by law.
- Paid parental leave, discounted meals, and exclusive perks across the Wonder family of brands.
- Benefit eligibility, effective dates, and plan options vary by employment classification and location.
Categories
Site Reliability
About Grubhub
Grubhub builds an online marketplace and logistics network for ordering takeout and delivery from local restaurants across the U.S. It earns commissions and delivery fees, and offers consumer subscriptions and corporate solutions. Founded in 2004 and headquartered in Chicago, it is part of Wonder Group and lists hundreds of thousands of merchants across 4,000+ cities.
