
Database Reliability Engineer - DBRE
Cognite - AI for IndustryResponsibilities
- Automate lifecycle management, provisioning, configuration, patching, upgrades, backups, and operational workflows for more than 1,000 PostgreSQL instances across Azure, AWS, and GCP.
- Operate and scale managed PostgreSQL services including Azure Database for PostgreSQL – Flexible Server, Amazon RDS, and GCP Cloud SQL.
- Design and operate Elasticsearch clusters across Elastic Cloud and ECK, including sizing, capacity planning, shard management, upgrades, backups, monitoring, disaster recovery, and security.
- Own Kafka cluster reliability and performance across self-managed Kubernetes-based and managed environments, including capacity planning, partition and topic design, replication, upgrades, backups, and schema evolution.
- Build Terraform- and Kubernetes-driven automation, operators, monitoring, alerting, and self-service infrastructure to reduce operational toil.
- Partner with Software Engineering, SRE, Platform Engineering, and Product teams to improve database reliability, scalability, security, resilience, and incident response.
Requirements
- At least 6 years of experience in Database Reliability Engineering, Database Engineering, SRE, Platform Engineering, or a closely related role.
- Strong hands-on experience operating PostgreSQL at scale, preferably in cloud-managed environments.
- Strong Elasticsearch experience including cluster administration, performance tuning, scaling, shard management, and troubleshooting.
- Experience operating databases and stateful workloads on Kubernetes.
- Experience with Kafka or another distributed database or storage system is valuable.
- Strong Infrastructure as Code experience with Terraform or similar tools.
- Experience with Azure, AWS, or GCP, with multi-cloud exposure preferred.
- Strong understanding of high availability, disaster recovery, backups, replication, capacity planning, and observability.
- Proficiency in Go, Python, or a similar programming or scripting language; Go proficiency and Python or another scripting language are preferred.
- Experience automating operational tasks, building self-service infrastructure, and troubleshooting complex distributed-system failures.
- Preferred experience includes Kafka, fdb-kubernetes-operator, Elastic Cloud, ECK, Kubernetes operators, database performance monitoring, private-cloud environments, and distributed-systems design.
Tech Stack
About Cognite - AI for Industry
Cognite is the only Industrial Data and AI platform built for scaling. We help Energy, Manufacturing, and Power & Utilities companies break down data silos and turn complex operational data into actionable, enterprise-wide value. Why Cognite? Unify: Cognite Data Fusion® unifies OT, ET, and IT data into a real-time Industrial Knowledge Graph. Automate: Cognite Atlas AI™ deploys low-code Industrial AI Agents to supercharge SME capacity. Deliver(2025): We drive a 400% ROI (Forrester) and have identified over $1B in customer value in the past year alone. Our Moonshot We are on a mission to deliver $100 billion in realized customer value by 2035. From increasing production uptime to reducing downtime by 30%, Cognite provides the foundation to scale AI solutions across every asset and site. Founded in 2016, we now number 700 strong, including some of the best software developers, data scientists, designers, and 3D specialists in the field. Get to know us better: Make data do more: makedatadomore.cognite.com/ Twitter: @CogniteData Facebook: @CogniteData IG: @CogniteData Youtube: /c/cognite