D.B. Group

Site Reliability Engineer, AVP

D.B. Group
Apply
15 hours ago
Bengaluru, IndiaStaff+

Responsibilities

  • Define and improve SLIs, SLOs, alerting standards, and error budgets for the CaaS Private platform.
  • Build and maintain metrics, logs, traces, alerts, and dashboards for platform health, capacity, latency, saturation, and failure analysis.
  • Investigate and resolve complex production issues across Kubernetes, Linux, networking, storage, ingress, service mesh, node services, platform dependencies, and tenant workloads.
  • Lead or coordinate incident response, stakeholder communication, blameless post-incident reviews, and durable follow-up actions.
  • Automate operational tasks, diagnostics, and remediation workflows to reduce toil and accelerate recovery.
  • Improve reliability, upgrade safety, resilience, capacity planning, performance, disaster recovery, and operational readiness.
  • Partner with platform, network, security, storage, and application teams on releases, troubleshooting, change execution, documentation, and SRE adoption.
  • Support the global GDC/Kubernetes platform in a 24x7 follow-the-sun model with on-call, weekend, public-holiday, and early Monday rotations.

Requirements

  • Bachelor’s degree in a technical or engineering discipline.
  • 10–12 years of hands-on experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure role.
  • Strong hands-on Kubernetes expertise covering cluster operations, upgrades, troubleshooting, networking, storage, security, Helm, Operators, and workload lifecycle management in bare-metal or private-cloud environments.
  • Experience with SLI/SLO/SLA management, error budgets, incident reduction, capacity planning, performance optimization, resilience engineering, and operational excellence.
  • Strong Linux system administration and networking fundamentals, including troubleshooting of complex infrastructure and distributed-system failures.
  • Experience operating observability solutions using Prometheus, Grafana, Splunk, distributed tracing, logging, alerting, dashboards, and OpenTelemetry-style concepts.
  • Strong automation and infrastructure scripting experience with Python, Bash, and Ansible.
  • Experience with CI/CD and GitOps platforms such as Argo CD, Jenkins, GitHub or Bitbucket, and Artifactory.
  • Understanding of incident management, root-cause analysis, operational readiness, virtualization, containerization, and distributed-system behavior under failure.
  • CKA or CKAD certification, or equivalent demonstrable Kubernetes expertise, is helpful.
  • Experience with Istio or Envoy, service-mesh observability, traffic management, Cilium, OPA Gatekeeper, admission controls, policy guardrails, container-native storage, load testing, chaos testing, failure injection, self-healing automation, and disaster-recovery validation is helpful.
  • Familiarity with PostgreSQL, Kafka, MongoDB, regulated or low-latency environments, and the ability to read Golang code are helpful.
  • Ability to use AI tools responsibly and communicate, collaborate, prioritize, and independently resolve complex problems.

Benefits

  • Best-in-class leave policy and gender-neutral parental leave.
  • 100% reimbursement under the gender-neutral childcare assistance benefit.
  • Sponsorship for industry-relevant certifications and education.
  • Employee Assistance Program for employees and family members.
  • Comprehensive hospitalization insurance for employees and dependents.
  • Accident and term life insurance.
  • Complementary health screening for employees aged 35 and above.
  • Training, career development, coaching, expert support, and a culture of continuous learning.
  • Flexible benefits that can be tailored to individual needs.
  • Bangalore, India location with a 24x7 follow-the-sun operating model, including on-call, weekend, public-holiday, and early Monday coverage as required.

Tech Stack

AmbassadorAnsibleApache KafkaArgo CDBashGoGrafanaHelmIstioJenkinsKubernetesLinuxMongoDBPostgreSQLPrometheusPythonSplunkTerraform

Categories

DevOpsSite Reliability
D.B. Group

About D.B. Group

501-1,000 employees

D.B. Group is an Italian freight forwarder and logistics provider offering air, ocean, road, and rail transport, customs brokerage, and supply chain management for global shippers. Founded in 1980 and headquartered in Montebelluna, Veneto, it is privately held and operates through 79 subsidiaries with a worldwide partner network. The company delivers end-to-end forwarding and integrated logistics services, from consolidation and warehousing to door-to-door deliveries and trade compliance.

Contact me