SolarWinds

Site Reliability Engineer

SolarWinds
Apply
20 hours ago
Bengaluru, IndiaSenior

Responsibilities

  • Operate, maintain, upgrade, and improve production Kubernetes clusters and workloads across AWS and Azure.
  • Manage Kubernetes platform components including Helm, Kustomize, operators, Istio, autoscaling, and cluster/node lifecycle management.
  • Support production database platforms such as ClickHouse and Aurora, including performance troubleshooting and operational health.
  • Build and maintain infrastructure using Terraform.
  • Develop automation and tooling with Python, Go, Bash, or similar technologies to reduce operational toil.
  • Participate in scheduled on-call rotations, respond to production incidents, and lead or contribute to incident resolution and root-cause analysis.
  • Improve observability, monitoring, logging, alerting, incident response, resilience, disaster recovery, and reliability practices.
  • Partner with software engineering and platform teams to design and deploy reliable, scalable services.
  • Contribute to capacity planning, performance optimization, patching, upgrades, infrastructure lifecycle management, documentation, runbooks, and troubleshooting guides.

Requirements

  • At least 2 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Platform Engineering, or a related field.
  • Strong hands-on production experience operating, upgrading, and troubleshooting Kubernetes clusters and workloads.
  • Strong understanding of Kubernetes fundamentals, networking, storage, scheduling, autoscaling, high availability patterns, upgrades, and cluster lifecycle management.
  • Strong hands-on experience with AWS and Azure cloud infrastructure.
  • Strong Linux systems administration and troubleshooting skills.
  • Experience operating customer-facing, highly available production systems and participating in production on-call rotations.
  • Strong Terraform and Infrastructure as Code experience.
  • Experience with scripting and automation using Python, Go, Bash, or similar languages.
  • Understanding of compute, networking, storage, DNS, load balancing, and security infrastructure concepts.
  • Strong troubleshooting, communication, collaboration, operational discipline, and documentation skills.

Tech Stack

Categories

Site Reliability
SolarWinds

About SolarWinds

1,001-5,000 employees

SolarWinds builds IT operations and observability software used by IT departments and MSPs to monitor networks, servers, databases, applications, logs, and help desks. It sells subscriptions and licenses for on-premises and SaaS products such as SolarWinds Observability and the Orion Platform; founded in 1999 and headquartered in Austin, Texas, the company is publicly traded on the NYSE under the ticker SWI.

Contact me