Overview
Senior Database Reliability Engineer (DBRE) for Cognite’s Cloud deployment team, owning reliability, scalability, automation, and operational excellence of core database infrastructure.
What you'll do
- Orchestrate and automate a 1000+ PostgreSQL fleet lifecycle across Azure, AWS, and GCP.
- Automate PostgreSQL provisioning, configuration, patching, upgrades, and backups.
- Operate and scale Elasticsearch clusters across Elastic Cloud and ECK, ensuring performance and reliability.
- Own Kafka cluster reliability and performance for high-throughput, low-latency event streaming.
- Establish best practices for sizing, capacity planning, shard management, upgrades, backups, and disaster recovery.
- Build monitoring and alerting for consumer lag, broker health, ISR status, and throughput.
What you'll need
- 6+ years of experience in Database Reliability Engineering, Database Engineering, SRE, Platform Engineering, or a related role.
- Hands-on experience operating PostgreSQL at scale, preferably in cloud-managed environments.
- Experience with Elasticsearch cluster administration, performance tuning, scaling, shard management, and troubleshooting.
- Experience operating databases/stateful workloads on Kubernetes.
- Infrastructure as Code experience with Terraform or similar tools.
- Proficiency scripting/programming in Python, Go, or a similar language.
Details
- Location: India (Bengaluru).
Read the full description and apply on the company’s own careers page.