4 months ago
Tel Aviv-Yafo, IsraelSenior
Responsibilities
- Own the reliability, availability, and performance of production environments.
- Lead efforts and develop long-term solutions to improve production stability and resiliency.
- Collaborate with development teams on architecture, system design, observability, and scalability.
- Perform capacity planning, system tuning, and infrastructure cost/performance optimization.
- Maintain and improve GitOps-driven deployment pipelines.
Requirements
- Have 7–8+ years of hands-on experience in SRE or DevOps roles managing production environments at scale.
- Demonstrate deep experience with AWS architecture, security, and best practices.
- Have strong programming skills in Python, Go, or TypeScript.
- Have at least 3 years of Infrastructure as Code experience with AWS CDK or Terraform.
- Have production experience with Kubernetes, including EKS, and microservices architecture.
- Have hands-on experience with CI/CD and GitOps tools such as GitHub Actions and ArgoCD.
- Understand Linux internals and networking fundamentals.
- Communicate effectively in English and Hebrew.
- Experience with observability stacks such as Datadog, Coralogix, Prometheus, or Grafana is preferred.
- Experience with Crossplane or advanced Helm usage is preferred.
- Experience working in a SaaS company is preferred.
Benefits
- Collaborative culture emphasizing product ownership and architectural planning.
- Opportunities to build résumé and portfolio experience.
- Competitive salary with opportunities for growth and advancement.
- On-site role with the Infrastructure team in Tel Aviv.
Tech Stack
Categories
DevOpsSite Reliability
