20 hours ago
Vila Nova de Gaia, PortugalSenior
Responsibilities
- Ensure the availability, performance, reliability, and resilience of cloud-native platforms and customer-facing services.
- Own observability through metrics, logs, traces, dashboards, alerting, and defined SLIs, SLOs, and SLAs.
- Lead incident response, troubleshooting, root cause analysis, and post-incident reviews.
- Identify bottlenecks and implement reliability, scalability, automation, and resilience improvements.
- Support CI/CD pipelines, infrastructure automation, deployment processes, and production environments.
- Design and operate distributed systems and participate in on-call rotations.
Requirements
- Strong understanding of Site Reliability Engineering principles, production operations, and reliability practices.
- Experience supporting high-availability production environments, including on-call support.
- Proven incident management, troubleshooting, and root cause analysis skills in complex distributed systems.
- Experience with Linux-based environments and system administration.
- Hands-on experience with Docker and Kubernetes.
- Knowledge of Prometheus, Grafana, ELK Stack, and Splunk for observability and monitoring.
- Experience building and maintaining CI/CD pipelines with GitHub Actions and Azure DevOps.
- Familiarity with Terraform and Ansible for infrastructure as code and configuration management.
- Understanding of networking fundamentals, cloud-native architectures, distributed systems, and incident, problem, and change management processes.
- Experience collaborating with development, infrastructure, and security teams in Agile environments.
Benefits
- Flexible hybrid work arrangement with flexibility to manage working hours and work remotely or from the office.
- Career Acceleration Programs supporting growth, reskilling, and new skills development.
- Empowering work environment with autonomy and peer collaboration.
- Health and life insurance.
- Referral program with bonuses and other partnership-based fringe benefits.
Tech Stack
Categories
Site Reliability
About Capgemini
Capgemini is a global IT services and consulting firm that delivers strategy, cloud, AI, software engineering, and managed services to large enterprises and public-sector clients. Founded in 1967 and headquartered in Paris, it is publicly traded on Euronext Paris and operates in 50+ countries. The group expanded its engineering capabilities by acquiring Altran in 2020, now operating as Capgemini Engineering.
