3 months ago
Madrid, SpainSenior
Responsibilities
- Implement and manage AWS cloud-native systems using automation and best-in-class tools.
- Operate and enhance Kubernetes clusters, deployment pipelines, service meshes, and infrastructure supporting a multi-tenant SaaS platform.
- Define and maintain SLOs, SLAs, and error budgets while addressing availability and performance issues.
- Develop infrastructure-as-code with Terraform or similar tools and build internal platform automation.
- Monitor infrastructure and applications, automate health checks and alerting, and improve operational efficiency.
- Participate in on-call rotations, troubleshoot outages, act as Incident Commander, and coordinate cross-team incident responses.
- Drive incident response improvements, reducing Mean Time to Detect and Mean Time to Recovery.
- Embed observability, fault tolerance, reliability, automated testing, canary deployments, and rollback strategies into service design.
- Contribute to security best practices, compliance automation, chaos engineering, resilience testing, and cost optimization.
Requirements
- Bachelor’s degree in Computer Science or equivalent practical experience.
- At least 5 years of experience as a Site Reliability Engineer or Platform Engineer.
- Strong software development, programming or scripting, public cloud, SaaS, and infrastructure-as-code experience.
- Proficiency with Kubernetes, Docker, Helm, CI/CD tools, monitoring solutions, and production systems.
- Experience with multi-tenant microservices architectures, incident management, on-call operations, and post-incident reviews.
- Deep understanding of Linux systems, networking, cloud architectures, service meshes, storage, and troubleshooting practices.
- Knowledge of zero-downtime, blue/green, and canary deployment strategies.
- Exposure to SOC 2, ISO 27001, HIPAA, or FedRAMP compliance standards; FedRAMP experience is preferred.
- Experience with chaos engineering or resilience testing is preferred.
- Strong problem-solving, communication, collaboration, organization, and English-language skills.
- Applicants are encouraged to apply even if they do not meet every listed requirement.
Benefits
- Permanent contract with a competitive compensation package.
- Hybrid work model balancing office and remote work, with a structured onboarding approach for new hires.
- Flexible hours and unlimited paid time off, plus 22 holidays, company-paid bank holidays, sick days, bereavement leave, and three annual volunteering days.
- Health insurance including outpatient dental, vision, health check-up, consultation, and pharmacy coverage.
- Access to professional training platforms and personal accident insurance.
- Maternity, paternity, and adoptive-parent leave benefits.
- Gratuity under the Payment of Gratuity Act after a minimum of five years of employment.
- Employee referral bonuses after successful hires complete three months of employment.
Tech Stack
AWSAzureBashDatadogDockerGitHub ActionsGitLab CI/CDGoGoogle Cloud PlatformHelmIstioJenkinsKubernetesLinuxPythonTerraform
Categories
Site Reliability
