2 hours ago
Paris, FranceSenior / Staff+
Responsibilities
- Monitor, operate, and maintain platform health, availability, performance, and scalability.
- Apply SRE principles and contribute to SLO, SLI, and Error Budget definition, tracking, and reporting.
- Participate in incident response, root cause analysis, post-incident reviews, and preventive and corrective actions.
- Develop automation for deployment, monitoring, recovery, and repetitive operational workflows.
- Implement and improve monitoring, logging, alerting, distributed tracing, and telemetry.
- Participate in resilience testing, chaos engineering, failover improvements, and disaster recovery activities.
- Collaborate with development and platform teams on reliability, operability, scalability, performance, change management, and operational readiness.
- Document operational procedures, runbooks, best practices, incidents, and improvements.
Requirements
- 7–10 years of professional experience in SRE, operations, DevOps, or platform engineering roles.
- Fluent written and spoken English and French.
- Strong communication and collaboration skills in a technical environment.
- Ability to work autonomously while contributing to a cross-functional team.
- A proactive, curious, solution-driven mindset with a strong focus on reliability, automation, and continuous improvement.
- Comfort operating in a fast-paced and evolving digital environment and constructively challenging existing practices.
Benefits
- Hybrid work arrangement with three days per week in Paris (8th arrondissement) after the trial period.
- 75% reimbursement of the monthly or annual public transport pass.
- Swile restaurant ticket card.
- Free access to a company-reserved gym.
- Sustainable mobility package.
- Health insurance and welfare benefits.
- Employee savings plan, profit sharing, and bonus.
- International, inclusive, and multidisciplinary work environment.
Categories
Site Reliability
