about 3 hours ago
Responsibilities
- Design and implement scalable, reliable systems across hybrid or multi-cloud environments.
- Drive improvements in system uptime, latency, and service health metrics.
- Build and manage infrastructure automation using Terraform, Helm, and Kubernetes.
- Enhance CI/CD pipelines for safe and automated rollouts.
- Own and improve the observability stack and define SLOs.
- Lead major incident response and root cause analysis.
- Ensure platform-level security and compliance with standards.
- Mentor junior SREs and contribute to technical roadmaps.
Requirements
- 5+ years of experience with Kubernetes, EKS, ECS, and containerized workloads.
- Expertise in AWS services such as EC2, Lambda, and RDS.
- Proficiency with Terraform, Helm, Jenkins, and GitOps.
- Deep understanding of observability frameworks and tools.
- Strong knowledge of Linux, networking fundamentals, and system performance tuning.
- Familiarity with Python, Go, or Shell scripting for automation.
- Practical experience in incident response and on-call operations.
Benefits
- Flexible work model with 2 days in the office and 3 days remote.
- Opportunities for career growth across multiple roles and locations.
- Collaborative and creative work environment.