7 months ago
Responsibilities
- Own and influence the incident management process end to end.
- Maintain and evolve the on-prem observability stack and alerting systems.
- Participate in the production on-call rotation to keep applications running smoothly.
- Develop automation and tools that support platform reliability.
- Contribute to production services with performance and resiliency in mind.
- Collaborate with product engineers to promote SRE principles across the R&D organization.
- Mentor SREs and product engineers and support reliability knowledge sharing.
Requirements
- Solid programming experience with Python, including Django and AsyncIO, and/or Java with Spring Boot.
- Experience maintaining an observability tools suite, specifically LGTM: Loki, Grafana, Tempo, and Mimir.
- Experience developing and maintaining Python services in production.
- Strong experience with AWS and Kubernetes.
- Proficiency with PostgreSQL and messaging systems such as RabbitMQ, NATS, and Kafka.
- Experienced on-call SRE engineer with hands-on distributed-systems troubleshooting experience.
- Proficiency in written and spoken English.
Benefits
- Remote-first approach with optional hybrid work from offices in Kyiv, Warsaw, and Lisbon.
- Long-term collaboration through employment contracts, employer-of-record, or B2B arrangements, with terms varying by location.
- Work schedule aligned with EU time zones.
- Professional and personal development within a collaborative, supportive team.
- Stable, growing SaaS product with ownership, startup energy, and strong technical challenges.
Tech Stack
Categories
DevOpsSite Reliability
About PandaDoc
Stand out with the top‑rated solution for creating, managing, tracking, and esigning every important document you handle.
