ArangoDB

Site Reliability Engineer (Europe)

ArangoDB
Apply
1 day ago
Remote, IndiaMid Level

Responsibilities

  • Design, implement, and maintain cloud infrastructure on AWS and Google Cloud.
  • Ensure the scalability, performance, and reliability of Kubernetes-based distributed database systems.
  • Write production-grade Golang code and automation for infrastructure management and system operations.
  • Optimize CI/CD pipelines, deployment processes, monitoring, logging, and alerting systems.
  • Develop disaster recovery, high-availability, and fault-tolerance strategies.
  • Identify bottlenecks and troubleshoot complex issues across networks, operating systems, and cloud infrastructure.
  • Participate in on-call rotations and respond to critical production incidents.
  • Collaborate with development, cross-functional, and Customer Success teams to improve reliability and resolve customer issues.

Requirements

  • Experience in an SRE or DevOps Engineer role in cloud-native environments.
  • Advanced knowledge of AWS or Google Cloud, Linux internals, Docker, and Kubernetes at scale.
  • Experience with Jenkins or CircleCI, Prometheus, Grafana, the ELK stack, and Git.
  • Knowledge of core networking, cloud security best practices, and systematic infrastructure troubleshooting.
  • Programming proficiency in Golang or Python.
  • Strong communication, autonomy, organization, and remote-team collaboration skills.
  • Experience with distributed databases or large-scale data storage systems is preferred.
  • Python or Bash scripting, Terraform, GitOps, and strong Golang automation experience are preferred.

Benefits

  • Remote position based in India.
  • Opportunity to work on AI and data infrastructure and help shape enterprise AI applications.
  • Collaboration with experienced engineering, marketing, and product teams.

Tech Stack

AWSBashCircleCIDockerGitGoGoogle CloudGoogle Cloud PlatformGrafanaJenkinsKubernetesLinuxPrometheusPythonTerraform

Categories

DevOpsSite Reliability
Contact me