6 days ago
Remote, Saudi ArabiaMid Level
Responsibilities
- Design and maintain highly available, fault-tolerant, and scalable infrastructure while eliminating single points of failure.
- Manage and optimize cloud workloads across AWS, GCP, or Azure, including cost and resource usage.
- Operate, troubleshoot, scale, and upgrade production Kubernetes clusters and containerized workloads.
- Implement monitoring and alerting with tools such as Prometheus, Grafana, Datadog, or ELK.
- Respond to incidents, lead root cause analysis, and improve reliability based on failure learnings.
- Write scripts and build automation to reduce manual operational work.
- Collaborate with DevOps and engineering teams on performance, CI/CD, deployment reliability, and reliability best practices.
Requirements
- Approximately three years of experience in SRE, DevOps, or infrastructure engineering.
- Hands-on experience with cloud environments such as AWS, GCP, or Azure and an understanding of distributed systems.
- Production experience with Kubernetes and troubleshooting Kubernetes issues.
- Experience using Terraform or similar infrastructure-as-code tools to manage infrastructure.
- Confidence working with Docker and Kubernetes and scripting with Python, Bash, or similar languages.
- Understanding of CI/CD pipelines and tools such as Jenkins, GitHub Actions, or Bitbucket.
- Knowledge of networking, load balancing, and high-availability design.
- Experience implementing monitoring and observability tools such as Prometheus, Grafana, Datadog, or ELK.
- Preferred experience with RabbitMQ or Redis, Ansible or AWX, multi-cloud or hybrid environments, cloud or Linux certifications, or ITI.
Benefits
- Remote work arrangement.
- Opportunity to own reliability and scalability improvements for a globally scaling AI-native customer experience platform.
Tech Stack
AnsibleAWSAzureBashDatadogDockerGitHub ActionsGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonRabbitMQRedisTerraform
Categories
Site Reliability
