
Database Reliability Engineer - DBRE
Cognite - AI for Industry2 months ago
Bengaluru, IndiaSenior
Responsibilities
- Standardize and automate the lifecycle management of more than 1,000 PostgreSQL instances across Azure, AWS, and GCP.
- Automate provisioning, configuration, patching, upgrades, backups, and operational workflows for managed PostgreSQL services.
- Design, operate, and scale Elasticsearch clusters across Elastic Cloud and ECK self-managed Kubernetes environments.
- Manage Elasticsearch cluster sizing, capacity planning, shard management, upgrades, backups, monitoring, disaster recovery, security, connectivity, and resilience.
- Operate Kafka clusters in self-managed Kubernetes-based and managed configurations across multi-cloud environments.
- Drive Kafka capacity planning, partition and topic design, replication strategy, broker performance tuning, upgrades, backups, disaster recovery, and schema evolution.
- Build monitoring and alerting for consumer lag, broker health, ISR status, throughput, and other reliability signals.
- Develop Terraform- and Kubernetes-driven automation, self-service infrastructure, and operational tooling.
- Partner with Software Engineering, SRE, Platform Engineering, and Product teams to scale database usage and improve reliability.
- Diagnose distributed-system failures, reduce operational toil, and improve availability, scalability, security, and observability.
Requirements
- At least 6 years of experience in Database Reliability Engineering, Database Engineering, SRE, Platform Engineering, or a closely related role.
- Strong hands-on experience operating PostgreSQL at scale, preferably in cloud-managed environments.
- Strong Elasticsearch experience including cluster administration, performance tuning, scaling, shard management, and troubleshooting.
- Experience operating databases and stateful workloads on Kubernetes.
- Experience with Kafka or another distributed database or storage system is valuable.
- Strong Infrastructure as Code experience with Terraform or similar tools.
- Experience with at least one major cloud platform—Azure, AWS, or GCP—with multi-cloud exposure preferred.
- Understanding of high availability, disaster recovery, backups, replication, capacity planning, and observability.
- Proficiency in Go, Python, or a similar programming or scripting language; Go proficiency and Python experience are preferred.
- Experience automating operational tasks and building self-service infrastructure.
- Strong troubleshooting and incident-management skills for complex distributed-system failures.
- Experience with Kafka, fdb-kubernetes-operator, Elastic Cloud, ECK, Kubernetes operators, stateful workload automation, observability platforms, database performance monitoring, distributed systems, data replication, consistency, and failure-domain design is advantageous.
Benefits
- Opportunity to work on large-scale distributed data infrastructure supporting a modern Knowledge Graph platform.
- Significant ownership across PostgreSQL, Elasticsearch, and Kafka infrastructure.
- International work environment with presence in Phoenix, Houston, Oslo, Tokyo, Bengaluru, and Abu Dhabi.
Tech Stack
Categories
DevOpsSite Reliability
About Cognite - AI for Industry
Cognite builds an industrial data and AI platform used by energy, manufacturing, and power & utilities companies to integrate OT/IT data and deploy AI at scale. Its flagship product, Cognite Data Fusion, is sold as enterprise SaaS with services for implementation and Industrial AI agents. Founded in 2016 and privately held, the company focuses on asset-intensive operations and Industrial DataOps.