
Senior Site Reliability Engineer / Kubernetes (Remote)
Pragmatike5 hours ago
Tallinn, Estonia or Bucharest, RomaniaSenior
Responsibilities
- Operate and maintain Debian/Ubuntu infrastructure and production Kubernetes clusters across bare-metal, virtualized, and on-premises environments.
- Manage Kubernetes upgrades, node pools, networking, storage, security hardening, and cluster lifecycle operations.
- Build automation and deployment workflows using Ansible, Bash, Python, GitOps, PXE boot, Preseed, and cloud-init.
- Design networking architecture covering VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Deploy and maintain observability stacks, optimize monitoring and alerting pipelines, and improve availability and latency.
- Define SLOs and SLIs across physical infrastructure, virtualization platforms, and software services.
- Lead incident response and escalations, maintain on-call coverage, and develop standard operating procedures.
- Manage OpenStack, Proxmox, and VMware virtualization and orchestration layers.
- Coordinate physical maintenance, hardware issues, and data-center operations for Policlouds.
- Support architecture, resource planning, quality improvements, and resource utilization across products while collaborating with engineering and cross-functional teams.
Requirements
- Expert-level hands-on experience operating Kubernetes in production environments.
- Strong network engineering skills, including VLANs, L2/L3 routing, VPNs, and multi-site connectivity.
- Strong proficiency with Linux systems administration, particularly Debian and Ubuntu.
- Experience designing complex network architectures and working with distributed systems and container orchestration.
- Experience building automation workflows with Ansible, Bash, Python, and Git-based practices.
- Experience with Prometheus, Grafana, ELK, Loki, or Graylog observability stacks.
- Experience with OpenStack, Proxmox, VMware, bare-metal provisioning, and MAAS.
- Experience with incident response, escalation procedures, on-call rotations, and reliability practices.
- Ability to develop SOPs and operational procedures, work autonomously, and collaborate effectively in a fast-paced engineering environment.
- Nice-to-have experience includes Istio, Linkerd, advanced CNI implementations, Cloudflare APIs, DNS automation, tunnel configurations, GPU infrastructure, RBAC, firewalls, network policies, IT asset management, license tracking, and multi-timezone coordination.
- Fluent English is mandatory.
Benefits
- Fully remote work within the EU timezone, CET ±2 hours, with flexible hours.
- ASAP start date.
- High-impact role with autonomy and ownership.
- Collaborative international engineering team.
- Work on cloud computing projects with a strong focus on reliability and automation.
Tech Stack
Categories
DevOpsSite Reliability
About Pragmatike
Pragmatike is a Paris-based IT services and recruiting firm connecting remote-first companies with software engineers and tech specialists worldwide. Founded in 2022, it places contractors or full-time hires and staffs teams to complete projects for startups and scaleups. The private partnership offers access to a large network of 50,000+ specialists across 60+ countries.