Nextiva

Staff Site Reliability Engineer

Nextiva
Apply
2 hours ago
Bengaluru, IndiaSenior / Staff+

Responsibilities

  • Own the reliability, availability, scalability, performance, and operational excellence of middleware and cloud infrastructure platforms.
  • Support and optimize Kafka environments through performance tuning, capacity planning, upgrades, and troubleshooting.
  • Administer, support, and optimize vector database platforms.
  • Manage and support GCP and GKE environments.
  • Drive infrastructure automation, platform reliability initiatives, and operational excellence.
  • Lead production incident response, troubleshooting, root-cause analysis, post-incident reviews, and escalation management.
  • Build and maintain monitoring, alerting, and observability solutions.
  • Define and maintain SLIs, SLOs, and error budgets.
  • Support database operations, performance optimization, and operational health.
  • Participate in a Follow the Sun support model and a weekend on-call rotation approximately once every four weeks.

Requirements

  • Senior-level candidates need 5+ years of relevant experience; staff-level candidates need 10+ years.
  • Strong experience operating distributed systems in production environments.
  • At least 2 years of hands-on production experience with vector databases.
  • Five or more years of production experience with GCP and Kubernetes, including production GKE management and troubleshooting.
  • Strong Kafka knowledge covering architecture, brokers, partitions, replication, consumer groups, monitoring, and troubleshooting; 5+ years of hands-on Kafka production experience is preferred.
  • Experience with incident management, root-cause analysis, infrastructure automation, CI/CD practices, and Infrastructure as Code; Terraform is preferred.
  • Automation and scripting experience using Python, Go, Shell, or similar technologies.
  • Experience with observability platforms such as Datadog, Splunk, OpenTelemetry, Prometheus, or Grafana.
  • Strong Linux administration and troubleshooting skills.
  • Experience supporting databases such as MongoDB, PostgreSQL, Redis, Elasticsearch, or ClickHouse.
  • Experience with Weaviate is strongly preferred.
  • Preferred qualifications include experience supporting AI-native platforms, LLM infrastructure, vector search, embedding technologies, and technical leadership or mentoring.
  • Practical use of AI tools such as ChatGPT and cloud AI services to improve troubleshooting, automation, and operational efficiency.

Benefits

  • Medical insurance coverage is available for employees, spouses, dependent children, and parents or in-laws, subject to stated coverage limits.
  • Group term and group personal accident insurance are provided, including stated coverage and insured-sum benefits.
  • The role provides 15 privilege-leave days, 6 paid sick-leave days, 6 casual-leave days, 26 weeks of paid maternity leave, 1 week of paternity leave, a birthday day off, and paid holidays.
  • Benefits include Provident Fund and Gratuity.
  • Employee Assistance Program and wellness initiatives are provided.
  • Employees receive ongoing learning and development opportunities and career advancement support.
  • This is an onsite Bengaluru role expected to work four days per week, with the possibility of five days as business needs require.
  • The role includes a Follow the Sun support model and weekend on-call approximately once every four weeks.

Tech Stack

Apache KafkaClickHouseDatadogElasticsearchGoGoogle Cloud PlatformGrafanaKubernetesLinuxMongoDBPostgreSQLPrometheusPythonRedisSplunkTerraform

Categories

DevOpsSite Reliability
Nextiva

About Nextiva

1,001-5,000 employees
Contact me