Nexthink

Senior Site Reliability Engineer

Nexthink
Apply
15 hours ago
Madrid, SpainSenior

Responsibilities

  • Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
  • Operate and enhance Kubernetes clusters, deployment pipelines, service meshes, and multi-tenant SaaS infrastructure.
  • Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
  • Develop infrastructure-as-code and internal platform automation for provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications, automate runbooks and alerting, and support reliable user experiences.
  • Participate in on-call rotations, troubleshoot outages, act as Incident Commander, and coordinate cross-team incident responses.
  • Drive incident response improvements, post-incident reviews, and reductions in MTTD and MTTR.
  • Work with software engineers to embed observability, fault tolerance, and reliability into service design.
  • Support automated testing, canary deployments, rollback strategies, compliance automation, security practices, and cost optimization.

Requirements

  • Bachelor's degree in Computer Science or equivalent practical experience.
  • 5+ years of experience as a Site Reliability Engineer or Platform Engineer.
  • Strong knowledge of software development best practices and programming or scripting with languages such as Python, Go, or Bash.
  • Hands-on experience with AWS, GCP, or Azure and supporting SaaS products.
  • Experience with Kubernetes, Docker, Helm, infrastructure-as-code, and CI/CD tools.
  • Experience supporting multi-tenant microservices architectures and managing production systems.
  • Experience with monitoring solutions such as Datadog and with rotating on-call schedules and critical incident management.
  • Deep understanding of Linux systems, networking, TCP/IP, VPNs, cloud architectures, service meshes, and storage.
  • Knowledge of zero-downtime, blue/green, and canary deployment strategies.
  • Exposure to SOC 2, ISO 27001, or HIPAA compliance standards; FedRAMP experience is a plus.
  • Experience with chaos engineering or resilience testing is preferred.
  • Strong troubleshooting, problem-solving, communication, collaboration, organization, and English-language skills.

Benefits

  • Permanent contract with a competitive compensation package including stock options.
  • Hybrid work model balancing office and remote work, with a structured approach for new hires.
  • Private Sanitas health insurance and company-covered daily meal vouchers of 11 EUR.
  • Flexible hours, unlimited paid time off in addition to 23 days of holidays, and three company-paid volunteer days.
  • Up to 25 EUR per month toward a gym subscription and flexible retribution for kindergarten and transport tickets.
  • Up to 50% reimbursement for English and Spanish classes.
  • Fresh fruit, snacks, company and team events, referral bonuses, and relocation support.
  • Base salary of €60,000–€85,500 gross per year; total on-target earnings are €66,000–€93,000 including an annual performance bonus.

Tech Stack

AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform

Categories

Site Reliability
Nexthink

About Nexthink

1,001-5,000 employees

Nexthink builds a digital employee experience (DEX) management platform used by enterprise IT teams to monitor endpoints, analyze performance, capture sentiment, and remediate issues at scale. It sells cloud-based software and services on a subscription model to improve reliability and support across the digital workplace. Founded in 2004 and headquartered in Prilly, Switzerland, the company is privately held and backed by Vista Equity Partners.

Contact me