Lucidya

Site Reliability Engineer

Lucidya
Apply
6 days ago
Remote, Saudi ArabiaMid Level

Responsibilities

  • Design and maintain highly available, fault-tolerant, and scalable infrastructure while eliminating single points of failure.
  • Manage and optimize cloud workloads across AWS, GCP, or Azure, including cost and resource usage.
  • Operate, troubleshoot, scale, and upgrade production Kubernetes clusters and containerized workloads.
  • Implement monitoring and alerting with tools such as Prometheus, Grafana, Datadog, or ELK.
  • Respond to incidents, lead root cause analysis, and improve reliability based on failure learnings.
  • Write scripts and build automation to reduce manual operational work.
  • Collaborate with DevOps and engineering teams on performance, CI/CD, deployment reliability, and reliability best practices.

Requirements

  • Approximately three years of experience in SRE, DevOps, or infrastructure engineering.
  • Hands-on experience with cloud environments such as AWS, GCP, or Azure and an understanding of distributed systems.
  • Production experience with Kubernetes and troubleshooting Kubernetes issues.
  • Experience using Terraform or similar infrastructure-as-code tools to manage infrastructure.
  • Confidence working with Docker and Kubernetes and scripting with Python, Bash, or similar languages.
  • Understanding of CI/CD pipelines and tools such as Jenkins, GitHub Actions, or Bitbucket.
  • Knowledge of networking, load balancing, and high-availability design.
  • Experience implementing monitoring and observability tools such as Prometheus, Grafana, Datadog, or ELK.
  • Preferred experience with RabbitMQ or Redis, Ansible or AWX, multi-cloud or hybrid environments, cloud or Linux certifications, or ITI.

Benefits

  • Remote work arrangement.
  • Opportunity to own reliability and scalability improvements for a globally scaling AI-native customer experience platform.

Tech Stack

AnsibleAWSAzureBashDatadogDockerGitHub ActionsGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonRabbitMQRedisTerraform

Categories

Site Reliability
Lucidya

About Lucidya

201-500 employees
Contact me