5 hours ago
Madrid, SpainSenior
Responsibilities
- Implement and manage cloud-native systems on AWS using automation and best-in-class tools.
- Operate and enhance Kubernetes clusters, deployment pipelines, and service meshes.
- Design, build, and maintain reliable, secure, and scalable infrastructure for a multi-tenant SaaS platform.
- Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
- Develop infrastructure-as-code for repeatable and auditable provisioning.
- Build internal platform tools and automation for provisioning, monitoring, and operational efficiency.
- Monitor infrastructure and applications to maintain high-quality user experiences.
- Participate in on-call rotations, respond to incidents, troubleshoot outages, and communicate resolutions.
- Act as Incident Commander and coordinate cross-team responses during critical incidents.
- Improve incident response processes and reduce Mean Time to Detect and Mean Time to Recovery.
- Diagnose and resolve complex production issues independently.
- Work with software engineers to embed observability, fault tolerance, and reliability into service design.
- Automate runbooks, health checks, and alerting.
- Support automated testing, canary deployments, and rollback strategies.
- Contribute to security best practices, compliance automation, resilience testing, and cost optimization.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
- Strong software development, programming, or scripting skills, including Python, Go, or Bash.
- Hands-on experience with public cloud services such as AWS, GCP, or Azure and support for SaaS products.
- Experience with Terraform or similar infrastructure-as-code tools.
- Proficiency with Kubernetes, Docker, Helm, and related container ecosystems.
- Experience supporting multi-tenant microservices architectures.
- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, FluxCD, or Crossplane.
- Experience managing monitoring solutions such as Datadog.
- Experience with production operations, rotating on-call schedules, critical incidents, and post-incident reviews.
- Strong system-level troubleshooting skills and knowledge of Linux, networking, cloud architectures, service meshes, and storage.
- Knowledge of TCP/IP, VPNs, VPCs, subnets, firewalls, load balancers, Istio, S3, and EBS.
- Knowledge of zero-downtime, blue/green, and canary deployment strategies.
- Exposure to SOC 2, ISO 27001, or HIPAA compliance standards; FedRAMP experience is a plus.
- Experience with chaos engineering or resilience testing is a plus.
- Strong problem-solving, communication, presentation, collaboration, organization, and English-language skills.
- The company welcomes applicants who do not meet every listed requirement and will assess candidates with different backgrounds and experience levels.
Benefits
- Permanent contract with a competitive compensation package.
- Hybrid work model balancing office and remote work, with structured onboarding for new hires.
- Flexible hours and unlimited paid time off in addition to 25 days of holidays.
- Three company-paid volunteer days.
- Free access to an on-site fitness center.
- Reimbursement of half-fare public transport travel cards.
- Reimbursement of up to 50% of French-language class costs.
- Fresh fruit, cookies, and soft drinks.
- Regular company and team events, including volunteer days, talks, team-building activities, and office meetups.
- Referral bonuses after successful hires complete three months of continuous employment.
- Relocation package for employees moving from another country.
- Benefits may vary for temporary, contract, and internship roles.
- The role is based in a hybrid office environment near Prilly-Malley train station.
Tech Stack
AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform
Categories
DevOpsSite Reliability
About Nexthink
Nexthink builds a digital employee experience (DEX) management platform used by enterprise IT teams to monitor endpoints, analyze performance, capture sentiment, and remediate issues at scale. It sells cloud-based software and services on a subscription model to improve reliability and support across the digital workplace. Founded in 2004 and headquartered in Prilly, Switzerland, the company is privately held and backed by Vista Equity Partners.
