9 days ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Build, scale, and maintain highly available, performant, and cost-efficient hybrid infrastructure across on-premises, public cloud, and AI/ML Kubernetes environments.
- Develop internal software tooling and manage infrastructure-as-code pipelines using Go, Python, or Rust to eliminate repetitive operations.
- Troubleshoot issues across CDN edge configurations, Linux kernel behavior, and network-layer bottlenecks.
- Design and maintain monitoring, telemetry, metrics collection, and alerting systems.
- Participate in on-call rotations, lead incident resolution, and conduct blameless post-mortems.
Requirements
- 7+ years of experience managing, scaling, and troubleshooting large-scale distributed Linux environments in production.
- Deep understanding of Linux system internals and TCP/IP, DNS, HTTP, and gRPC, with hands-on experience with Fastly, Cloudflare, Akamai, or CloudFront.
- Hands-on experience with infrastructure-as-code and orchestration tools such as Terraform, Ansible, Puppet, ArgoCD, or Jenkins.
- Production experience managing containerized environments with Kubernetes and Docker.
- Solid programming skills in at least one of Go, Python, or Rust.
- Experience designing and operating large-scale telemetry, metrics collection, and alerting stacks such as Prometheus, Grafana, and ELK/logging is a bonus.
- Practical experience optimizing infrastructure costs and resource efficiency across cloud and on-premises environments is a bonus.
Benefits
- Comprehensive benefits including health benefits, a fully stocked kitchen, gym partnerships, and parking-related perks.
- Hybrid work schedule with three days in the office and flexibility.
- Work with global publishers, clients, and technology partners.
Tech Stack
Categories
DevOpsSite Reliability
