
Platform Engineer (remote work)
CloudLinux1 day ago
Remote, Spain +9 moreSenior
Responsibilities
- Operate and maintain the observability platform, including onboarding, cost and capacity management, and alerting.
- Administer GitLab and the CI runner fleet, including upgrades, capacity, access, backups, and restore drills.
- Keep platform services healthy through monitoring, runbooks, documentation, and repeatable operations.
- Research, design, and deploy new services as code with monitoring, backups, and documentation.
- Support developers with access, onboarding, pipeline issues, exporters, and dashboards, and turn recurring requests into self-service.
- Lead incident diagnosis, mitigation, service restoration, root-cause analysis, post-mortems, and prevention improvements.
- Deliver infrastructure changes through reviewed merge requests and write technical runbooks, guides, notices, and status updates.
- Use AI engineering assistants to delegate scoped work, review outputs, and verify generated automation before production use.
Requirements
- Senior-level infrastructure, platform, or site reliability engineering experience with responsibility for at least one production service.
- Linux systems administration and debugging experience on bare metal and virtual machines.
- Production Kubernetes experience, including GitOps delivery and personally performed cluster upgrades.
- Infrastructure-as-code experience using Ansible and Terraform or OpenTofu with merge-request review workflows.
- Production GitLab administration and GitLab CI experience, or equivalent depth with another CI system.
- Working knowledge of Prometheus and Grafana, including operating them for a team, writing alert rules and dashboards, and reading PromQL.
- Ability to write technical explanations, runbooks, notices, and request responses for engineers outside the team.
- Strong communication and interpersonal skills for scoping, prioritizing, coordinating, and communicating platform work with product teams.
- Advanced use of AI engineering assistants such as Claude and Codex, including context provision, task decomposition, agent-loop design, debugging, testing, and output verification.
- Upper-intermediate or higher English proficiency.
- Preferred experience with SLOs, burn-rate alerts, data-based thresholds, Kata Containers, Firecracker, gVisor, Ceph RGW, AWS cost work, Sentry, Kafka, ClickHouse, Redis, Python, or Go.
Benefits
- Fully remote work with flexible working hours from any location worldwide.
- Professional development, interesting projects, and an education budget.
- 24 paid vacation days, 10 national holidays, and unlimited sick leave.
- Compensation for private medical insurance.
- Coworking and gym or sports reimbursement.
- Opportunity to receive a reward for an innovative idea that the company can patent.
Tech Stack
Categories
DevOpsSite Reliability
About CloudLinux
CloudLinux builds a commercially supported Linux operating system and server-hardening tools for hosting providers and data centers running multi-tenant environments. It sells subscriptions and support that improve server stability, resource isolation, and security for shared hosting workloads, along with related maintenance and patching services. Founded in 2009 and headquartered in Estero, Florida, the company is privately held.