Pragmatike

Senior Site Reliability Engineer (Remote)

Pragmatike
Apply
5 hours ago
Bucharest, RomaniaSenior

Responsibilities

  • Operate and maintain Debian/Ubuntu Linux infrastructure.
  • Deploy, manage, scale, upgrade, and secure Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
  • Build automation and deployment workflows using Ansible, Bash/Python, PXE boot, Preseed, cloud-init, and Git-based workflows.
  • Design and maintain networking architecture, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Deploy and maintain observability stacks using Prometheus, Grafana, Loki, ELK, and Graylog.
  • Lead incident response and escalation, improve availability and latency, and define SLOs and SLIs across infrastructure layers.
  • Optimize alerting and monitoring pipelines and establish on-call schedules across time zones.
  • Develop standard operating procedures for recurring operations and maintenance.
  • Coordinate physical maintenance, hardware issues, and data-center operations for Policlouds.
  • Manage OpenStack, Proxmox, and VMware virtualization and orchestration layers.
  • Contribute to product architecture, resource planning, system quality, and resource utilization improvements.
  • Collaborate with development teams and cross-functional stakeholders, including Hivenet, Policloud, and Customer Success teams.

Requirements

  • Expert-level hands-on experience operating Kubernetes in production.
  • Strong network engineering skills covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
  • Strong Linux systems administration experience with Debian and Ubuntu.
  • Experience designing complex network architectures and operating distributed systems and container orchestration.
  • Experience building automation workflows with Ansible, Bash/Python, and Git-based tools.
  • Experience with observability stacks such as Prometheus, Grafana, ELK, Loki, or Graylog.
  • Experience with OpenStack, Proxmox, or VMware virtualization technologies.
  • Experience with bare-metal provisioning and MAAS.
  • Experience with incident response, escalation procedures, and on-call rotations.
  • Experience creating standard operating procedures and operational frameworks from scratch.
  • Fluent English and the ability to work autonomously in a fast-paced engineering environment.
  • Preferred experience with Istio, Linkerd, advanced CNI implementations, Cloudflare APIs, DNS automation, tunnel configurations, GPU infrastructure, security practices, IT asset management, distributed teams, and SRE frameworks.

Benefits

  • Fully remote work within the EU timezone, CET ±2 hours, with flexible hours.
  • ASAP start date.
  • High-impact role with autonomy and ownership.
  • Collaborative international engineering team.
  • Reliability- and automation-focused technology environment.

Tech Stack

AnsibleBashCloudflareGrafanaGraylogIstioKubernetesLinuxOpenStackPrometheusPython

Categories

DevOpsSite Reliability
Pragmatike

About Pragmatike

1-10 employees

Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.

Contact me