5 months ago
Bengaluru, IndiaSenior
Responsibilities
- Design, build, and operate highly available, scalable clusters for MongoDB, Elasticsearch, and Apache Kafka.
- Own major platform components and deliver complex initiatives from design through production.
- Develop architectural patterns and platform standards for reliability, scalability, and performance.
- Troubleshoot complex distributed-systems issues across multi-cluster environments.
- Build and maintain infrastructure as code and deployment automation.
- Improve observability, reliability, and operational excellence across the data platform.
- Collaborate with engineering teams and Product Management to align platform capabilities with product needs.
- Mentor junior engineers and contribute to knowledge sharing.
- Participate in on-call rotations and continuous platform improvement.
Requirements
- At least 5 years of experience in platform engineering, SRE, or infrastructure-focused roles.
- Production experience operating at least one of MongoDB, Apache Kafka, or Elasticsearch, including day-2 operations.
- Experience designing and operating Kubernetes environments; Helm and ArgoCD experience is a significant plus.
- Proficiency in Python, Java, or a similar programming language.
- Experience with infrastructure as code tools such as Terraform.
- Hands-on experience with major cloud platforms, preferably AWS.
- Experience implementing and maintaining observability solutions such as Prometheus/Grafana or ELK.
- Understanding of security principles and experience embedding security best practices into production environments and code promotion systems.
- Strong collaboration, communication, problem-solving, and attention-to-detail skills.
- Experience with internal platform APIs or automation tooling, regulated environments, or self-service infrastructure is desirable.
