Nexthink

Senior Site Reliability Engineer

Nexthink
Apply
5 hours ago
Madrid, SpainSenior

Responsibilities

  • Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
  • Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes.
  • Design, build, and maintain reliable, secure, and scalable infrastructure for a multi-tenant SaaS platform.
  • Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
  • Develop infrastructure-as-code for repeatable and auditable provisioning.
  • Build internal platform tools and automation for provisioning, monitoring, and operational efficiency.
  • Monitor infrastructure and applications to maintain high-quality user experiences.
  • Participate in on-call rotations, respond to incidents, troubleshoot outages, and communicate resolutions.
  • Act as Incident Commander and coordinate cross-team responses during critical incidents.
  • Improve incident response processes and reduce Mean Time to Detect and Mean Time to Recovery.
  • Diagnose and resolve complex production issues independently.
  • Work with software engineers to embed observability, fault tolerance, and reliability into service design.
  • Automate runbooks, health checks, and alerting.
  • Support automated testing, canary deployments, and rollback strategies.
  • Contribute to security best practices, compliance automation, resilience testing, and cost optimization.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
  • Strong software development, programming, or scripting skills, including Python, Go, or Bash.
  • Hands-on experience with public cloud services such as AWS, GCP, or Azure and support for SaaS products.
  • Experience with Terraform or similar infrastructure-as-code tools.
  • Proficiency with Kubernetes, Docker, Helm, and related container ecosystems.
  • Experience supporting multi-tenant microservices architectures.
  • Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, FluxCD, or Crossplane.
  • Experience managing monitoring solutions such as Datadog.
  • Experience with production operations, rotating on-call schedules, critical incidents, and post-incident reviews.
  • Strong system-level troubleshooting skills and knowledge of Linux, networking, cloud architectures, service meshes, and storage.
  • Knowledge of TCP/IP, VPNs, VPCs, subnets, firewalls, load balancers, Istio, S3, and EBS.
  • Knowledge of zero-downtime, blue/green, and canary deployment strategies.
  • Exposure to SOC 2, ISO 27001, or HIPAA compliance standards; FedRAMP experience is a plus.
  • Experience with chaos engineering or resilience testing is a plus.
  • Strong problem-solving, communication, presentation, collaboration, organization, and English-language skills.
  • The company welcomes applicants who do not meet every listed requirement and will assess candidates with different backgrounds and experience levels.

Benefits

  • Permanent contract with a competitive compensation package.
  • Hybrid work model balancing office and remote work, with structured onboarding for new hires.
  • Flexible hours and unlimited paid time off in addition to 25 days of holidays.
  • Three company-paid volunteer days.
  • Free access to an on-site fitness center.
  • Reimbursement of half-fare public transport travel cards.
  • Reimbursement of up to 50% of French-language class costs.
  • Fresh fruit, cookies, and soft drinks.
  • Regular company and team events, including volunteer days, talks, team-building activities, and office meetups.
  • Referral bonuses after successful hires complete three months of continuous employment.
  • Relocation package for employees moving from another country.
  • Benefits may vary for temporary, contract, and internship roles.
  • The role is based in a hybrid office environment near Prilly-Malley train station.

Tech Stack

AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform

Categories

DevOpsSite Reliability
Nexthink

About Nexthink

1,001-5,000 employees

Nexthink builds a digital employee experience (DEX) management platform used by enterprise IT teams to monitor endpoints, analyze performance, capture sentiment, and remediate issues at scale. It sells cloud-based software and services on a subscription model to improve reliability and support across the digital workplace. Founded in 2004 and headquartered in Prilly, Switzerland, the company is privately held and backed by Vista Equity Partners.

Contact me