Site Reliability Engineer - IDP
Intermedia Intelligent Communications4 months ago
Remote, PortugalSenior
Responsibilities
- Ensure the availability, performance, and reliability of critical applications and services through monitoring, alerting, and optimization.
- Define and maintain SLIs, SLOs, error budgets, reliability standards, and operational guardrails.
- Partner with development and platform teams to improve application performance, resilience, and self-service capabilities.
- Automate deployments, rollbacks, scaling, failover, recovery, and other operational tasks.
- Improve CI/CD pipelines and integrate automated validation, reliability checks, and progressive delivery practices.
- Build and maintain observability capabilities including metrics, logs, traces, dashboards, alerts, and operational views.
- Respond to production incidents, troubleshoot issues, conduct root cause analyses, and drive corrective actions.
- Run fire drills, game days, and chaos engineering exercises to validate resilience.
- Monitor resource usage, capacity trends, and scaling behavior.
- Partner with security teams on secure communication, access controls, and data protection.
- Lead or contribute to production readiness, operational review, incident review, and reliability improvement activities.
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Proven experience in Site Reliability Engineering, Platform Engineering, or Infrastructure/DevOps roles with strong operational ownership.
- Strong expertise in application monitoring, observability platforms, incident response, and production troubleshooting.
- Strong understanding of SLIs, SLOs, error budgets, alerting quality, and incident management.
- Proficiency in scripting and automation using Python, Bash, Terraform, Ansible, or similar tools and languages.
- Experience with AWS, Azure, or Google Cloud.
- Strong knowledge of CI/CD pipelines, deployment automation, progressive delivery, infrastructure as code, and configuration management.
- Experience with Docker and Kubernetes.
- Strong problem-solving skills, operational judgment, attention to detail, communication, and cross-team collaboration.
- Preferred experience with chaos engineering, internal platforms, developer portals, golden paths, service catalogs, observability stacks, telemetry standards, microservices, multi-tenant environments, UCaaS or CCaaS platforms, capacity planning, resilience testing, and operational readiness.
Benefits
- Primarily remote work with occasional visits to the Coimbra office; future offices are planned in Aveiro and Porto.