3 months ago
Paris, FranceMid Level / Senior
Responsibilities
- Create, maintain, and improve observability and alerting tools and frameworks for software engineers.
- Own the Service Level Objectives framework and support the design and maintenance of Service Level Indicators and objectives.
- Define and improve incident-management practices, standards, post-mortems, and chaos-engineering processes.
- Act as Incident Commander during high-severity incidents when needed.
- Develop and maintain reliability automation tools, including Terraform modules and Go applications.
- Build reporting on operational metrics and incidents to support continuous improvement.
- Use AI and agentic approaches to reduce toil and streamline daily work.
Requirements
- 3–7 years of experience in SRE, DevOps, or software engineering roles.
- Strong knowledge of observability tools and metrics, logging, and tracing.
- Production troubleshooting and on-call experience diagnosing and resolving technical issues; Kubernetes experience is a plus.
- Full working proficiency in English.
- Strong communication skills and the ability to work across multidisciplinary teams.
- Ability to take ownership while aligning work with business priorities and adapting to different contexts.
- Familiarity with incident-management platforms such as Grafana IRM is preferred.
- Experience with SLOs and SLIs, OpenTelemetry integration, or programming in Go is preferred.
- Familiarity with object-oriented programming, scripting languages, and web or mobile testing tools is advantageous.
Benefits
- Hybrid work arrangement with 2–3 days in the office.
- Four additional weeks beyond legal maternity and paternity leave.
- 50% healthcare coverage through Alan.
- Financial support for home-office equipment.
- At least 25 days of holiday per year.
- Local meal-plan policy through a Swile card.
- 50% transportation reimbursement through Forfait Mobilité Durable.
- Free unlimited carpooling and bus rides.
- Training, mentorship, and internal mobility opportunities.
- Employee stock ownership plan and regular team-building events.
- One paid day off per year to test the BlaBlaCar product.
- Hiring process typically lasts 25–30 days, includes video interviews and a system-design interview, and has one onsite interview.
Tech Stack
Categories
Site Reliability
