
Senior Site Reliability Engineer (MAAS)
Pragmatike10 hours ago
Bucharest, RomaniaSenior
Responsibilities
- Operate and maintain large-scale Debian/Ubuntu Linux infrastructure across bare-metal and virtualized environments.
- Own MAAS-based provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
- Operate production Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting.
- Design and maintain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.
- Automate infrastructure provisioning and operations with Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
- Build and maintain deployment workflows using PXE, Preseed, and cloud-init.
- Operate observability platforms and improve SLIs, SLOs, alerting, and reliability practices.
- Lead infrastructure incident response, troubleshooting, escalation, post-incident improvements, and on-call operations.
- Manage hardware-layer infrastructure including IPMI/Redfish, BMCs, RAID, storage, diagnostics, and GPU systems.
- Manage virtualization platforms and GPU passthrough where required.
- Build internal tooling for host discovery, configuration, IPAM, hardware health, and operational automation.
- Own infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and runbooks.
- Collaborate with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency.
Requirements
- 5+ years of hands-on SRE, infrastructure, systems, or platform engineering experience.
- Expert Linux administration experience, particularly with Debian and Ubuntu.
- Strong production experience with MAAS and bare-metal provisioning.
- Expert hands-on experience operating Kubernetes in production, including lifecycle management, networking, storage, upgrades, and troubleshooting.
- Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
- Strong automation skills with Ansible, Bash, and/or Python.
- Experience with Terraform or OpenTofu and Git-based infrastructure workflows.
- Production experience with Prometheus, Grafana, or comparable observability platforms.
- Experience with incident response, on-call operations, monitoring, alerting, and reliability practices.
- Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies.
- Experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
- Strong understanding of distributed systems, container orchestration, and infrastructure reliability.
- Experience with infrastructure security, including RBAC, firewalls, network policies, secrets management, and security hardening.
- Ability to create SOPs, runbooks, and operational processes from scratch.
- Fluent English and the ability to work autonomously in a fast-paced engineering environment.
- Preferred experience includes GPU infrastructure, Proxmox VE with ZFS/Ceph and GPU passthrough, VictoriaMetrics/VictoriaLogs, NetBox, Vault, SOPS, Atlantis, Cloudflare APIs, UniFi, service mesh, advanced CNI implementations, Go, Ceph, distributed databases, multi-site operations, and establishing SRE frameworks.
Benefits
- Fully remote role for candidates working in the EU timezone, CET ±2 hours.
- Flexible working hours.
- High-impact position with significant technical ownership and autonomy.
- Opportunity to work directly with bare-metal, Kubernetes, networking, and GPU infrastructure.
- International, engineering-driven team with a strong focus on automation, reliability, and infrastructure at scale.
- Opportunity to shape the architecture and operational foundations of a growing cloud platform.
- Start date is ASAP.
Categories
DevOpsSite Reliability
About Pragmatike
Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.