Pragmatike

Senior Site Reliability Engineer (MAAS)

Pragmatike
Apply
10 hours ago
Bucharest, RomaniaSenior

Responsibilities

  • Operate and maintain large-scale Debian/Ubuntu Linux infrastructure across bare-metal and virtualized environments.
  • Own MAAS-based provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
  • Operate production Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting.
  • Design and maintain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.
  • Automate infrastructure provisioning and operations with Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
  • Build and maintain deployment workflows using PXE, Preseed, and cloud-init.
  • Operate observability platforms and improve SLIs, SLOs, alerting, and reliability practices.
  • Lead infrastructure incident response, troubleshooting, escalation, post-incident improvements, and on-call operations.
  • Manage hardware-layer infrastructure including IPMI/Redfish, BMCs, RAID, storage, diagnostics, and GPU systems.
  • Manage virtualization platforms and GPU passthrough where required.
  • Build internal tooling for host discovery, configuration, IPAM, hardware health, and operational automation.
  • Own infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and runbooks.
  • Collaborate with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency.

Requirements

  • 5+ years of hands-on SRE, infrastructure, systems, or platform engineering experience.
  • Expert Linux administration experience, particularly with Debian and Ubuntu.
  • Strong production experience with MAAS and bare-metal provisioning.
  • Expert hands-on experience operating Kubernetes in production, including lifecycle management, networking, storage, upgrades, and troubleshooting.
  • Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
  • Strong automation skills with Ansible, Bash, and/or Python.
  • Experience with Terraform or OpenTofu and Git-based infrastructure workflows.
  • Production experience with Prometheus, Grafana, or comparable observability platforms.
  • Experience with incident response, on-call operations, monitoring, alerting, and reliability practices.
  • Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies.
  • Experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
  • Strong understanding of distributed systems, container orchestration, and infrastructure reliability.
  • Experience with infrastructure security, including RBAC, firewalls, network policies, secrets management, and security hardening.
  • Ability to create SOPs, runbooks, and operational processes from scratch.
  • Fluent English and the ability to work autonomously in a fast-paced engineering environment.
  • Preferred experience includes GPU infrastructure, Proxmox VE with ZFS/Ceph and GPU passthrough, VictoriaMetrics/VictoriaLogs, NetBox, Vault, SOPS, Atlantis, Cloudflare APIs, UniFi, service mesh, advanced CNI implementations, Go, Ceph, distributed databases, multi-site operations, and establishing SRE frameworks.

Benefits

  • Fully remote role for candidates working in the EU timezone, CET ±2 hours.
  • Flexible working hours.
  • High-impact position with significant technical ownership and autonomy.
  • Opportunity to work directly with bare-metal, Kubernetes, networking, and GPU infrastructure.
  • International, engineering-driven team with a strong focus on automation, reliability, and infrastructure at scale.
  • Opportunity to shape the architecture and operational foundations of a growing cloud platform.
  • Start date is ASAP.

Tech Stack

AnsibleBashCloudflareGitGoGrafanaKubernetesLinuxOpenStackPrometheusPythonTerraformVault

Categories

DevOpsSite Reliability
Pragmatike

About Pragmatike

1-10 employees

Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.

Contact me