6 months ago
Responsibilities
- Own and influence the incident-management process end-to-end.
- Maintain and evolve the on-prem observability stack and alerting systems.
- Participate in the production on-call rotation and keep applications running smoothly.
- Develop automation and tools that support platform reliability.
- Contribute to production services with performance and resiliency in mind.
- Collaborate with product engineers to promote SRE principles across the R&D organization.
- Mentor SRE team members or product engineers.
Requirements
- Solid programming experience with Python, including Django and AsyncIO, and/or Java with Spring Boot.
- Experience maintaining an observability tools suite, specifically LGTM: Loki, Grafana, Tempo, and Mimir.
- Experience developing and maintaining Python services in production.
- Strong experience with AWS and Kubernetes.
- Proficiency with PostgreSQL and messaging systems such as RabbitMQ, NATS, and Kafka.
- Experience as an on-call SRE engineer and troubleshooting distributed systems in production environments.
- Proficiency in written and spoken English.
Benefits
- Remote-first work with optional hybrid work from offices in Kyiv, Warsaw, and Lisbon.
- Long-term collaboration through employment contracts, employer-of-record, or B2B arrangements, with contract terms and benefits varying by location.
- Work schedule aligned with EU time zones.
- Professional and personal development in a collaborative and supportive team.
- Stable, growing SaaS product with ownership, startup energy, and challenging technical work.
Tech Stack
About PandaDoc
Stand out with the top‑rated solution for creating, managing, tracking, and esigning every important document you handle.
