15 hours ago
Madrid, SpainSenior
Responsibilities
- Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
- Operate and enhance Kubernetes clusters, deployment pipelines, service meshes, and multi-tenant SaaS infrastructure.
- Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
- Develop infrastructure-as-code and internal platform automation for provisioning, monitoring, and operational efficiency.
- Monitor infrastructure and applications, automate runbooks and alerting, and support reliable user experiences.
- Participate in on-call rotations, troubleshoot outages, act as Incident Commander, and coordinate cross-team incident responses.
- Drive incident response improvements, post-incident reviews, and reductions in MTTD and MTTR.
- Work with software engineers to embed observability, fault tolerance, and reliability into service design.
- Support automated testing, canary deployments, rollback strategies, compliance automation, security practices, and cost optimization.
Requirements
- Bachelor's degree in Computer Science or equivalent practical experience.
- 5+ years of experience as a Site Reliability Engineer or Platform Engineer.
- Strong knowledge of software development best practices and programming or scripting with languages such as Python, Go, or Bash.
- Hands-on experience with AWS, GCP, or Azure and supporting SaaS products.
- Experience with Kubernetes, Docker, Helm, infrastructure-as-code, and CI/CD tools.
- Experience supporting multi-tenant microservices architectures and managing production systems.
- Experience with monitoring solutions such as Datadog and with rotating on-call schedules and critical incident management.
- Deep understanding of Linux systems, networking, TCP/IP, VPNs, cloud architectures, service meshes, and storage.
- Knowledge of zero-downtime, blue/green, and canary deployment strategies.
- Exposure to SOC 2, ISO 27001, or HIPAA compliance standards; FedRAMP experience is a plus.
- Experience with chaos engineering or resilience testing is preferred.
- Strong troubleshooting, problem-solving, communication, collaboration, organization, and English-language skills.
Benefits
- Permanent contract with a competitive compensation package including stock options.
- Hybrid work model balancing office and remote work, with a structured approach for new hires.
- Private Sanitas health insurance and company-covered daily meal vouchers of 11 EUR.
- Flexible hours, unlimited paid time off in addition to 23 days of holidays, and three company-paid volunteer days.
- Up to 25 EUR per month toward a gym subscription and flexible retribution for kindergarten and transport tickets.
- Up to 50% reimbursement for English and Spanish classes.
- Fresh fruit, snacks, company and team events, referral bonuses, and relocation support.
- Base salary of €60,000–€85,500 gross per year; total on-target earnings are €66,000–€93,000 including an annual performance bonus.
Tech Stack
AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform
Categories
Site Reliability
About Nexthink
Nexthink builds a digital employee experience (DEX) management platform used by enterprise IT teams to monitor endpoints, analyze performance, capture sentiment, and remediate issues at scale. It sells cloud-based software and services on a subscription model to improve reliability and support across the digital workplace. Founded in 2004 and headquartered in Prilly, Switzerland, the company is privately held and backed by Vista Equity Partners.
