
Senior Site Reliability Engineer
SolarWinds5 hours ago
Kraków, PolandSenior
Responsibilities
- Operate, maintain, upgrade, and improve production Kubernetes clusters and workloads across AWS and Azure.
- Manage Kubernetes components and technologies including Helm, Kustomize, operators, Istio, autoscaling, and cluster and node lifecycle management.
- Support production database platforms including ClickHouse, Aurora, and other distributed data systems, including performance troubleshooting and operational health.
- Build and maintain infrastructure using Terraform.
- Develop automation and tooling with Python, Go, Bash, or similar technologies to reduce operational toil and improve reliability.
- Participate in scheduled on-call rotations, respond to production incidents, and lead or contribute to incident resolution and root-cause analysis.
- Improve observability, monitoring, logging, alerting, and incident response across infrastructure and services.
- Partner with software engineering and platform teams to design and deploy reliable, scalable services.
- Contribute to capacity planning, performance optimization, patching, upgrades, infrastructure lifecycle management, disaster recovery, and resilience initiatives.
- Develop and maintain operational documentation, runbooks, and troubleshooting guides.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, Platform Engineering, or a related field.
- Strong hands-on production experience operating, upgrading, and troubleshooting Kubernetes clusters and workloads.
- Strong understanding of Kubernetes fundamentals, including workloads, networking, storage, scheduling, autoscaling, high-availability patterns, and cluster lifecycle management.
- Strong hands-on experience with AWS and Azure cloud infrastructure.
- Strong Linux systems administration and troubleshooting skills.
- Experience operating customer-facing, highly available production systems and participating in production on-call rotations.
- Strong experience with Terraform and Infrastructure as Code.
- Experience with scripting and automation using Python, Go, Bash, or similar languages.
- Strong understanding of infrastructure concepts including compute, networking, storage, DNS, load balancing, and security.
- Strong troubleshooting, communication, collaboration, documentation, and incident-response skills.
- Ability to stay calm during incidents, take ownership, learn continuously, and drive reliability-focused improvements.
Benefits
- Hybrid 3+2 work arrangement with at least three office days and two home-office days; Wednesdays and Thursdays are mandatory office days.
- Full-time employment contract in Kraków, Poland.
- 10 study days and 2 volunteering days per year.
- 30-day holidays after five years of tenure and sabbatical leave.
- Four weeks of paternity leave.
- Up to 8700 PLN per year for personal education.
- Medical care through Luxmed with individual, partner, or family packages fully paid by the company.
- Company-paid group life insurance and a pension scheme with a 1.5% employer contribution.
- Unlimited LinkedIn Learning access and English/Polish classes.
- MyBenefit platform subsidy, available vouchers and Multisport cards, race fee reimbursement, employee assistance, referral and appreciation programs, and free office lunches on Wednesdays.
Categories
DevOpsSite Reliability