Firmus Technologies

Senior Database Reliability Engineer

Firmus Technologies
Apply
8 hours ago
Singapore, SingaporeSenior

Responsibilities

  • Define database architecture, scaling, high-availability, failover, tenant-isolation, connection-pooling, proxy, and multi-site standards.
  • Own database provisioning, configuration, upgrades, backups, restores, decommissioning, and self-service workflows through infrastructure as code and automation.
  • Drive production database reliability through SLOs, capacity planning, query tuning, monitoring, incident response, on-call, post-mortems, and operational runbooks.
  • Set RPO and RTO targets and regularly prove recovery through restore and failover drills.
  • Own database access control, encryption, audit logging, change control, and compliance evidence for SOC 2 Type 2 and ISO 27001.
  • Partner with software, platform, data engineering, and observability teams on schema design, migrations, data modelling, CDC, replicas, and telemetry.
  • Support customer escalations involving database reliability, performance, or recovery.
  • Occasionally travel overseas when required.

Requirements

  • Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
  • At least 7 years of database reliability engineering experience, including at least 3 years owning PostgreSQL in large-scale production.
  • Deep PostgreSQL troubleshooting experience involving planner behaviour, locking, bloat, connection storms, replication, failover, backup and restore, and query tuning.
  • Hands-on experience with PostgreSQL high-availability, backup, and connection-pooling tools such as Patroni-style HA, pgBackRest or WAL-G, and pgBouncer or similar.
  • Production ownership of Redis or Memcached, including high availability or clustering, memory and eviction management, persistence, and failover decisions.
  • Production experience with a document or NoSQL database such as MongoDB, Cassandra, or DynamoDB, including high availability, replication, backup and restore, and failure diagnosis.
  • Experience automating database operations with infrastructure as code and Python or Go; experience with stateful databases on Kubernetes and operators is valued.
  • Experience defining RPO and RTO, running recovery drills, and owning SLOs, on-call rotations, and post-mortems that produced lasting fixes.
  • Understanding of Linux storage and networking effects on database performance and experience with database monitoring tools such as Prometheus, Grafana, or OpenTelemetry.
  • Experience operating databases under compliance frameworks such as SOC 2 Type 2 or ISO 27001.
  • Willingness to participate in incident-response on-call and communicate clearly in written and spoken English.

Benefits

  • Full-time employment based in Singapore.
  • Occasional overseas travel may be required.
  • Inclusive workplace committed to diversity and sustainability.

Tech Stack

Amazon DynamoDBApache CassandraGoGrafanaKubernetesLinuxMongoDBPostgreSQLPrometheusPythonRedis

Categories

Firmus Technologies

About Firmus Technologies

51-200 employees

Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.

Contact me