SolarWinds

Senior Site Reliability Engineer

SolarWinds
Apply
5 hours ago
Kraków, PolandSenior

Responsibilities

  • Operate, maintain, upgrade, and improve production Kubernetes clusters and workloads across AWS and Azure.
  • Manage Kubernetes components and technologies including Helm, Kustomize, operators, Istio, autoscaling, and cluster and node lifecycle management.
  • Support production database platforms including ClickHouse, Aurora, and other distributed data systems, including performance troubleshooting and operational health.
  • Build and maintain infrastructure using Terraform.
  • Develop automation and tooling with Python, Go, Bash, or similar technologies to reduce operational toil and improve reliability.
  • Participate in scheduled on-call rotations, respond to production incidents, and lead or contribute to incident resolution and root-cause analysis.
  • Improve observability, monitoring, logging, alerting, and incident response across infrastructure and services.
  • Partner with software engineering and platform teams to design and deploy reliable, scalable services.
  • Contribute to capacity planning, performance optimization, patching, upgrades, infrastructure lifecycle management, disaster recovery, and resilience initiatives.
  • Develop and maintain operational documentation, runbooks, and troubleshooting guides.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Platform Engineering, or a related field.
  • Strong hands-on production experience operating, upgrading, and troubleshooting Kubernetes clusters and workloads.
  • Strong understanding of Kubernetes fundamentals, including workloads, networking, storage, scheduling, autoscaling, high-availability patterns, and cluster lifecycle management.
  • Strong hands-on experience with AWS and Azure cloud infrastructure.
  • Strong Linux systems administration and troubleshooting skills.
  • Experience operating customer-facing, highly available production systems and participating in production on-call rotations.
  • Strong experience with Terraform and Infrastructure as Code.
  • Experience with scripting and automation using Python, Go, Bash, or similar languages.
  • Strong understanding of infrastructure concepts including compute, networking, storage, DNS, load balancing, and security.
  • Strong troubleshooting, communication, collaboration, documentation, and incident-response skills.
  • Ability to stay calm during incidents, take ownership, learn continuously, and drive reliability-focused improvements.

Benefits

  • Hybrid 3+2 work arrangement with at least three office days and two home-office days; Wednesdays and Thursdays are mandatory office days.
  • Full-time employment contract in Kraków, Poland.
  • 10 study days and 2 volunteering days per year.
  • 30-day holidays after five years of tenure and sabbatical leave.
  • Four weeks of paternity leave.
  • Up to 8700 PLN per year for personal education.
  • Medical care through Luxmed with individual, partner, or family packages fully paid by the company.
  • Company-paid group life insurance and a pension scheme with a 1.5% employer contribution.
  • Unlimited LinkedIn Learning access and English/Polish classes.
  • MyBenefit platform subsidy, available vouchers and Multisport cards, race fee reimbursement, employee assistance, referral and appreciation programs, and free office lunches on Wednesdays.

Tech Stack

Categories

DevOpsSite Reliability
SolarWinds

About SolarWinds

1,001-5,000 employees
Contact me