22 days ago
Remote, EMEASenior
Responsibilities
- Define, measure, and improve SLOs, SLAs, and error budgets for critical services.
- Scale the platform to support growing customers and transaction volumes.
- Identify weak points, risky patterns, and capacity limits before they affect customers.
- Participate in a shared 24/7 on-call rotation and lead incident response, mitigation, blameless postmortems, and follow-through.
- Own observability through dashboards, alerting, tracing, and logging, and improve performance across latency-sensitive paths.
- Design, implement, test, and deliver new product features across Java/Spring Boot microservices.
- Collaborate with product and engineering teams to deliver reliable, well-tested services.
Requirements
- 5+ years of experience building and operating distributed backend services with Java and Spring Boot, AWS cloud services, Kafka, and PostgreSQL.
- Experience operating high-availability systems and scaling systems under increasing load.
- Fluency with modern observability practices and strong SLO/SLA knowledge.
- A proactive ownership mindset focused on preventing problems before they occur.
- Ability to own features end-to-end from design and implementation through testing and delivery.
- Proven hands-on experience using AI effectively as an engineering tool.
- Payments, card issuing, transaction processing, PCI DSS, fraud prevention, chaos or resilience testing, capacity planning, load testing, and canary, blue/green, or progressive deployment experience are preferred.
Benefits
- Attractive remuneration.
- Choice of preferred operating system: Windows or Mac.
- Remote-work flexibility for an EU/UK remote role.
- Pliant Card with monthly credit for exploring the product and enjoying food with colleagues.
- Opportunity to work in a growing team with knowledge sharing, professional development, and transparent communication.
Tech Stack
Categories
BackendSite Reliability
