6 months ago
Responsibilities
- Own and influence the incident management process end-to-end.
- Maintain and evolve the on-premises observability stack.
- Participate in the on-call rotation to keep production applications running smoothly.
- Develop automation and tools that support platform reliability.
- Contribute to production services with performance and resiliency in mind.
- Collaborate with product engineers to promote SRE principles across the R&D organization.
- Mentor SRE team members or product engineers.
Requirements
- Solid programming experience with Python, including Django and AsyncIO, and/or Java with Spring Boot.
- Experience maintaining an observability tools suite, specifically Loki, Grafana, Tempo, and Mimir.
- Experience developing and maintaining Python services in production.
- Strong experience with AWS and Kubernetes.
- Proficiency with PostgreSQL and messaging systems such as RabbitMQ, NATS, and Kafka.
- Experience working as an on-call SRE engineer and troubleshooting distributed systems in production.
- Proficiency in written and spoken English.
Benefits
- Remote-first work with optional hybrid work from offices in Kyiv, Warsaw, and Lisbon.
- Flexible collaboration through employment contracts, employer-of-record, or B2B arrangements, with terms varying by location.
- Work schedule aligned with EU time zones.
- Professional and personal development in a collaborative and supportive team.
- Stable, growing SaaS product with ownership, start-up energy, and challenging technical work.
Tech Stack
Categories
DevOpsSite Reliability
About PandaDoc
Stand out with the top‑rated solution for creating, managing, tracking, and esigning every important document you handle.
