Nexthink

Senior Site Reliability Engineer

Nexthink
Apply
3 months ago
Madrid, SpainSenior

Responsibilities

  • Implement and manage AWS cloud-native systems using automation and best-in-class tools.
  • Operate and enhance Kubernetes clusters, deployment pipelines, service meshes, and infrastructure supporting a multi-tenant SaaS platform.
  • Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
  • Develop infrastructure-as-code with Terraform or similar tools and build internal platform automation.
  • Monitor infrastructure and applications, automate health checks and alerting, and improve operational efficiency.
  • Participate in on-call rotations, troubleshoot outages, act as Incident Commander, and coordinate cross-team incident responses.
  • Drive incident response improvements, reducing Mean Time to Detect and Mean Time to Recovery.
  • Embed observability, fault tolerance, reliability, automated testing, canary deployments, and rollback strategies into service design.
  • Contribute to security best practices, compliance automation, chaos engineering, resilience testing, and cost optimization.

Requirements

  • Bachelor’s degree in Computer Science or equivalent practical experience.
  • At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
  • Strong software development, programming or scripting, public cloud, SaaS, and infrastructure-as-code experience.
  • Proficiency with Kubernetes, Docker, Helm, CI/CD tools, monitoring solutions, and production systems.
  • Experience with multi-tenant microservices architectures, incident management, on-call operations, and post-incident reviews.
  • Deep understanding of Linux systems, networking, cloud architectures, service meshes, storage, and troubleshooting practices.
  • Knowledge of zero-downtime, blue/green, and canary deployment strategies.
  • Exposure to SOC 2, ISO 27001, HIPAA, or FedRAMP compliance standards; FedRAMP experience is preferred.
  • Experience with chaos engineering or resilience testing is preferred.
  • Strong problem-solving, communication, collaboration, organization, and English-language skills.
  • Applicants are encouraged to apply even if they do not meet every listed requirement.

Benefits

  • Permanent contract with a competitive compensation package.
  • Hybrid work model balancing office and remote work, with a structured onboarding approach for new hires.
  • Flexible hours and unlimited paid time off, plus 22 holidays, company-paid bank holidays, sick days, bereavement leave, and three annual volunteering days.
  • Health insurance including outpatient dental, vision, health check-up, consultation, and pharmacy coverage.
  • Access to professional training platforms and personal accident insurance.
  • Maternity, paternity, and adoptive-parent leave benefits.
  • Gratuity under the Payment of Gratuity Act after a minimum of five years of employment.
  • Employee referral bonuses after successful hires complete three months of employment.

Tech Stack

AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform

Categories

Site Reliability
Nexthink

About Nexthink

1,001-5,000 employees
Contact me