2 hours ago
Bengaluru, IndiaSenior / Staff+
Responsibilities
- Own the reliability, availability, scalability, performance, and operational excellence of middleware and cloud infrastructure platforms.
- Support and optimize Kafka environments through performance tuning, capacity planning, upgrades, and troubleshooting.
- Administer, support, and optimize vector database platforms.
- Manage and support GCP and GKE environments.
- Drive infrastructure automation, platform reliability initiatives, and operational excellence.
- Lead production incident response, troubleshooting, root-cause analysis, post-incident reviews, and escalation management.
- Build and maintain monitoring, alerting, and observability solutions.
- Define and maintain SLIs, SLOs, and error budgets.
- Support database operations, performance optimization, and operational health.
- Participate in a Follow the Sun support model and a weekend on-call rotation approximately once every four weeks.
Requirements
- Senior-level candidates need 5+ years of relevant experience; staff-level candidates need 10+ years.
- Strong experience operating distributed systems in production environments.
- At least 2 years of hands-on production experience with vector databases.
- Five or more years of production experience with GCP and Kubernetes, including production GKE management and troubleshooting.
- Strong Kafka knowledge covering architecture, brokers, partitions, replication, consumer groups, monitoring, and troubleshooting; 5+ years of hands-on Kafka production experience is preferred.
- Experience with incident management, root-cause analysis, infrastructure automation, CI/CD practices, and Infrastructure as Code; Terraform is preferred.
- Automation and scripting experience using Python, Go, Shell, or similar technologies.
- Experience with observability platforms such as Datadog, Splunk, OpenTelemetry, Prometheus, or Grafana.
- Strong Linux administration and troubleshooting skills.
- Experience supporting databases such as MongoDB, PostgreSQL, Redis, Elasticsearch, or ClickHouse.
- Experience with Weaviate is strongly preferred.
- Preferred qualifications include experience supporting AI-native platforms, LLM infrastructure, vector search, embedding technologies, and technical leadership or mentoring.
- Practical use of AI tools such as ChatGPT and cloud AI services to improve troubleshooting, automation, and operational efficiency.
Benefits
- Medical insurance coverage is available for employees, spouses, dependent children, and parents or in-laws, subject to stated coverage limits.
- Group term and group personal accident insurance are provided, including stated coverage and insured-sum benefits.
- The role provides 15 privilege-leave days, 6 paid sick-leave days, 6 casual-leave days, 26 weeks of paid maternity leave, 1 week of paternity leave, a birthday day off, and paid holidays.
- Benefits include Provident Fund and Gratuity.
- Employee Assistance Program and wellness initiatives are provided.
- Employees receive ongoing learning and development opportunities and career advancement support.
- This is an onsite Bengaluru role expected to work four days per week, with the possibility of five days as business needs require.
- The role includes a Follow the Sun support model and weekend on-call approximately once every four weeks.
Tech Stack
Apache KafkaClickHouseDatadogElasticsearchGoGoogle Cloud PlatformGrafanaKubernetesLinuxMongoDBPostgreSQLPrometheusPythonRedisSplunkTerraform
Categories
DevOpsSite Reliability
