CloudLinux

Senior Database Reliability Engineer (DBRE) (remote work)

CloudLinux
Apply
2 months ago
Remote, EMEA +6 moreSenior

Responsibilities

  • Own production PostgreSQL reliability, including HA design, Patroni, PgBouncer, replication, failover, upgrades, vacuum and bloat control, query tuning, locks, indexes, capacity, backups, PITR, and restore validation.
  • Improve disaster recovery through tested restores, documented recovery paths, measurable RTO/RPO targets, runbooks, and safe maintenance plans.
  • Support ClickHouse, MongoDB, and Redis by troubleshooting incidents, reviewing access and data-safety changes, improving monitoring, and developing ClickHouse operational expertise.
  • Automate provisioning, grants, backups, restores, health checks, and ownership metadata using Ansible, Terraform/OpenTofu, GitLab CI/CD, scripts, and reproducible runbooks.
  • Help build DBaaS-style self-service capabilities for database requests, access, credentials, and operational checks.
  • Improve observability and incident response using Grafana, metrics, logs, SLOs, alert rules, Opsgenie routing, and clear production-incident communication.

Requirements

  • Deep hands-on PostgreSQL experience in business-critical production environments, typically 5+ years or equivalent depth.
  • Strong knowledge of PostgreSQL internals and operations, including MVCC, WAL, transactions, locks, indexes, query planning, replication, autovacuum, bloat, major upgrades, backups, PITR, and restore testing.
  • Experience with highly available databases and the ability to reason about quorum, split-brain risk, failover, rollback, and recovery.
  • Strong Linux and infrastructure fundamentals covering systemd, networking, storage, filesystems, CPU, memory, disk bottlenecks, TLS, DNS, firewalls, and root-cause troubleshooting.
  • Automation skills with Ansible and scripting; Terraform/OpenTofu and GitLab CI/CD experience are strong advantages.
  • Ability and willingness to support multiple database engines and learn ClickHouse quickly.
  • Practical use of AI engineering assistants such as Claude and Codex, with personal verification of generated SQL, commands, scripts, and operational conclusions.
  • Upper-intermediate or higher English proficiency for clear team communication.
  • Preferred experience with ClickHouse operations, Keeper/ZooKeeper, MergeTree engines, distributed DDL, grants, row policies, backups, query troubleshooting, and cluster recovery.
  • Preferred experience with MongoDB replica sets, Percona Backup for MongoDB, Redis/Sentinel, database observability, SLOs, alert tuning, incident runbooks, internal platforms, self-service portals, or DBaaS workflows.

Benefits

  • Fully remote work with flexible working hours from any location worldwide.
  • 24 paid vacation days per year, 10 national holidays, and unlimited sick leave.
  • Compensation for private medical insurance.
  • Co-working and gym/sports reimbursement.
  • Education budget.
  • Opportunity to receive a reward for an innovative idea that the company can patent.
  • Professional development and challenging production-infrastructure projects.

Tech Stack

AnsibleClickHouseGitLab CI/CDGrafanaLinuxMongoDBPostgreSQLRedisTerraform

Categories

Site Reliability
CloudLinux

About CloudLinux

201-500 employees
Contact me