over 3 years ago
Paris, France or Berlin, GermanySenior
Responsibilities
- Operate and safeguard PostgreSQL, Redis, Kafka, and Elasticsearch in production.
- Lead incident response and root-cause analysis for storage, query, and performance problems.
- Define and prioritize backup and disaster-recovery scenarios with the Risk & Compliance team.
- Build internal platform APIs and self-service tooling for backend teams to provision storage resources.
- Help make storage APIs safe and traceable for autonomous AI agents.
- Own reliability and disaster-recovery roadmap work in a regulated fintech environment.
Requirements
- Production experience operating PostgreSQL, including performance investigation, backups, restores, and safe changes.
- Working knowledge of Redis, Kafka, or Elasticsearch, with the ability to learn additional systems quickly.
- Experience handling production incidents involving stateful systems.
- Ability to ship platform features and code, not only scripts.
- Practical daily use of AI for debugging, coding, documentation, or operational work.
- Interest in enabling safe and traceable AI-agent access to sensitive production systems.
Benefits
- Fully remote-first work with teammates distributed across Europe.
- Real ownership from day one over reliability and disaster-recovery initiatives.
- Opportunity to help shape resilience practices for a regulated, banking-licensed fintech.
- Join a growing Storage team expected to scale to approximately 5–6 engineers.
- Hiring process averages 20 working days.
Tech Stack
Categories
Site Reliability
