
Senior Site Reliability Engineer / Kubernetes (Remote)
Pragmatike4 hours ago
Rome, Italy +4 moreSenior
Responsibilities
- Operate and maintain Debian/Ubuntu-based Linux infrastructure.
- Deploy, manage, scale, upgrade, and secure Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
- Design and maintain VLAN, L2/L3 routing, VPN, and multi-site networking architecture.
- Automate provisioning and operations using Ansible, Bash, Python, GitOps workflows, PXE boot, Preseed, and cloud-init.
- Deploy and maintain observability stacks using Prometheus, Grafana, Loki, ELK, and Graylog.
- Lead incident response and escalations, improve availability and latency, and optimize monitoring and alerting pipelines.
- Define SLOs and SLIs, establish on-call schedules, and develop standard operating procedures.
- Coordinate physical maintenance, manage OpenStack, Proxmox, and VMware layers, and contribute to overall product architecture.
- Plan infrastructure resources, improve resource utilization with development teams, and collaborate with Hivenet, Policloud, and Customer Success teams.
Requirements
- Expert-level hands-on experience operating Kubernetes in production.
- Strong network engineering skills covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Strong Linux systems administration skills with Debian and Ubuntu.
- Experience building automation workflows with Ansible, Bash, Python, and Git-based tools.
- Experience with Prometheus, Grafana, ELK, Loki, or Graylog observability stacks.
- Experience with OpenStack, Proxmox, or VMware virtualization technologies.
- Experience with bare-metal provisioning and MAAS.
- Understanding of distributed systems and container orchestration.
- Experience with incident response, escalation procedures, on-call rotations, and operational procedure development.
- Fluent English and the ability to work autonomously in an engineering-driven environment.
- Preferred experience with Istio, Linkerd, advanced CNI implementations, Cloudflare APIs, DNS automation, tunnel configurations, GPU infrastructure, security practices, IT asset management, multi-timezone teams, and SRE frameworks.
Benefits
- Fully remote work in the EU timezone, CET ±2 hours, with flexible hours.
- ASAP start date.
- High-impact role with autonomy and ownership.
- Collaborative international engineering team and reliability- and automation-focused technology stack.
Tech Stack
Categories
DevOpsSite Reliability
About Pragmatike
Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.