Senior Database Reliability Engineer
Firmus Technologies8 hours ago
Singapore, SingaporeSenior
Responsibilities
- Define database architecture, scaling, high-availability, failover, tenant-isolation, connection-pooling, proxy, and multi-site standards.
- Own database provisioning, configuration, upgrades, backups, restores, decommissioning, and self-service workflows through infrastructure as code and automation.
- Drive production database reliability through SLOs, capacity planning, query tuning, monitoring, incident response, on-call, post-mortems, and operational runbooks.
- Set RPO and RTO targets and regularly prove recovery through restore and failover drills.
- Own database access control, encryption, audit logging, change control, and compliance evidence for SOC 2 Type 2 and ISO 27001.
- Partner with software, platform, data engineering, and observability teams on schema design, migrations, data modelling, CDC, replicas, and telemetry.
- Support customer escalations involving database reliability, performance, or recovery.
- Occasionally travel overseas when required.
Requirements
- Bachelor's degree in computer science or a related technical field, or equivalent practical experience.
- At least 7 years of database reliability engineering experience, including at least 3 years owning PostgreSQL in large-scale production.
- Deep PostgreSQL troubleshooting experience involving planner behaviour, locking, bloat, connection storms, replication, failover, backup and restore, and query tuning.
- Hands-on experience with PostgreSQL high-availability, backup, and connection-pooling tools such as Patroni-style HA, pgBackRest or WAL-G, and pgBouncer or similar.
- Production ownership of Redis or Memcached, including high availability or clustering, memory and eviction management, persistence, and failover decisions.
- Production experience with a document or NoSQL database such as MongoDB, Cassandra, or DynamoDB, including high availability, replication, backup and restore, and failure diagnosis.
- Experience automating database operations with infrastructure as code and Python or Go; experience with stateful databases on Kubernetes and operators is valued.
- Experience defining RPO and RTO, running recovery drills, and owning SLOs, on-call rotations, and post-mortems that produced lasting fixes.
- Understanding of Linux storage and networking effects on database performance and experience with database monitoring tools such as Prometheus, Grafana, or OpenTelemetry.
- Experience operating databases under compliance frameworks such as SOC 2 Type 2 or ISO 27001.
- Willingness to participate in incident-response on-call and communicate clearly in written and spoken English.
Benefits
- Full-time employment based in Singapore.
- Occasional overseas travel may be required.
- Inclusive workplace committed to diversity and sustainability.
Tech Stack
About Firmus Technologies
Firmus Technologies builds energy‑efficient AI infrastructure, developing liquid‑cooled “AI Factory” data centers and operating a large‑scale GPU cloud for model training. The company sells capacity and services to developers, enterprises, education, and government customers, with a focus on energy and cost efficiency across Asia‑Pacific. Founded in 2019 in Australia, Firmus is privately held and headquartered in St Leonards, Tasmania.