
Platform Engineer (remote work)
CloudLinux1 day ago
Remote, Spain +9 moreSenior
Responsibilities
- Operate and maintain the observability platform, including team onboarding, cost and capacity management, and alerting.
- Administer GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills.
- Keep production services healthy with monitoring, runbooks, and reliable operational practices.
- Research, design, and deploy new services as code with monitoring, backups, and documentation.
- Support developers with access, onboarding, pipeline issues, exporters, and dashboards while turning recurring requests into self-service.
- Respond to incidents, restore service safely, complete root-cause analyses and post-mortems, and implement prevention or detection improvements.
- Deliver infrastructure changes through reviewed merge requests and write runbooks, onboarding guides, maintenance notices, and status updates.
- Use AI engineering assistants to delegate scoped work, review generated output, and verify automation before production use.
Requirements
- Senior-level experience in infrastructure, platform, or site reliability engineering, including responsibility for at least one production service.
- Linux systems administration and debugging on bare metal and virtual machines.
- Production Kubernetes experience delivered through GitOps, including independently performing cluster upgrades.
- Infrastructure as code experience with Ansible and Terraform or OpenTofu, with changes reviewed in merge requests.
- Production GitLab administration and GitLab CI experience, or equivalent depth with another CI system.
- Working knowledge of Prometheus and Grafana, including operating them for a team, writing alert rules and dashboards, and reading PromQL.
- Ability to write clear technical explanations, runbooks, notices, and responses for engineers outside the team.
- Strong communication and interpersonal skills for scoping, prioritizing, coordinating, and communicating work with product teams.
- Advanced use of AI engineering assistants such as Claude and Codex, including context provision, task decomposition, agent loops, unattended execution, debugging, testing, and output verification.
- Upper-intermediate or higher English proficiency.
- Preferred experience with SLOs, burn-rate alerts, data-driven alert thresholds, Kata Containers, Firecracker, gVisor, Ceph RGW or similar S3-compatible storage, AWS cost work, self-hosted Sentry, Kafka, ClickHouse, Redis, Python, or Go.
Benefits
- Fully remote work with flexible working hours from anywhere worldwide.
- Professional development focus, interesting projects, and an education budget.
- 24 paid vacation days, 10 national holidays, and unlimited sick leave.
- Private medical insurance compensation.
- Co-working and gym or sports reimbursement.
- Opportunity to receive a reward for an innovative idea that the company can patent.
Tech Stack
Categories
DevOpsSite Reliability
About CloudLinux
CloudLinux builds a commercially supported Linux operating system and server-hardening tools for hosting providers and data centers running multi-tenant environments. It sells subscriptions and support that improve server stability, resource isolation, and security for shared hosting workloads, along with related maintenance and patching services. Founded in 2009 and headquartered in Estero, Florida, the company is privately held.