Pragmatike

Senior Site Reliability Engineer / Kubernetes (Remote)

Pragmatike
Apply
4 hours ago
Rome, Italy +4 moreSenior

Responsibilities

  • Operate and maintain Debian/Ubuntu-based Linux infrastructure.
  • Deploy, manage, scale, upgrade, and secure Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
  • Automate provisioning and operations using Ansible, Bash, Python, and GitOps workflows.
  • Design and maintain VLAN, L2/L3 routing, VPN, and multi-site networking architecture.
  • Build automated deployment workflows using PXE boot, Preseed, and cloud-init.
  • Deploy and maintain Prometheus/Grafana, Loki, ELK, and Graylog observability stacks.
  • Lead incident response and escalation, improve availability and latency, and optimize monitoring and alerting.
  • Define SLOs and SLIs across physical, virtualization, platform, and software-service layers.
  • Establish on-call schedules and develop standard operating procedures for recurring operations.
  • Coordinate physical maintenance, hardware issues, and data-center operations for Policlouds.
  • Manage OpenStack, Proxmox, and VMware virtualization and orchestration layers.
  • Contribute to product architecture, resource planning, and development-team initiatives to improve quality and resource utilization.
  • Collaborate with Hivenet, Policloud, Customer Success, and distributed cross-functional stakeholders.

Requirements

  • Expert-level hands-on experience operating Kubernetes in production.
  • Strong network engineering skills covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Strong proficiency in Linux systems administration, particularly Debian and Ubuntu.
  • Experience designing complex network architectures and working with distributed systems and container orchestration.
  • Experience building and maintaining automation workflows with Ansible, Bash/Python, and Git-based workflows.
  • Experience with observability technologies including Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience with OpenStack, Proxmox, VMware, bare-metal provisioning, and MAAS.
  • Experience with incident response, escalation procedures, on-call rotations, and reliability practices.
  • Ability to create SOPs and operational procedures from scratch and work autonomously in an engineering-driven environment.
  • Nice-to-have experience with Istio, Linkerd, advanced CNI implementations, Cloudflare APIs, DNS automation, tunnel configurations, GPU infrastructure, security practices, IT asset management, or distributed-team coordination.
  • Fluent English is mandatory.

Benefits

  • Fully remote work within the EU timezone, CET ±2 hours.
  • Flexible hours.
  • High-impact role with autonomy and ownership.
  • Collaborative international engineering team.
  • Opportunity to work with a technology stack focused on reliability and automation.
  • ASAP start date.

Tech Stack

AnsibleBashCloudflareGitGrafanaGraylogIstioKubernetesLinuxOpenStackPrometheusPython

Categories

DevOpsSite Reliability
Pragmatike

About Pragmatike

1-10 employees

Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.

Contact me