6 months ago
Responsibilities
- Own and influence the incident management process end to end
- Maintain and evolve the on-prem observability stack and alerting systems
- Participate in the on-call rotation to keep production applications running smoothly
- Develop automations and tools that support platform reliability
- Contribute to production services with performance and resiliency in mind
- Collaborate with product engineers to promote SRE principles across the R&D organization
- Mentor SRE team members or product engineers
Requirements
- Solid programming experience with Python, including Django and AsyncIO, and/or Java with Spring Boot
- Experience maintaining an observability tools suite, specifically LGTM: Loki, Grafana, Tempo, and Mimir
- Experience developing and maintaining Python services in production
- Strong experience with AWS and Kubernetes
- Proficiency with PostgreSQL and messaging systems such as RabbitMQ, NATS, and Kafka
- Experience as an on-call SRE engineer and troubleshooting distributed systems in production
- Proficiency in written and spoken English
Benefits
- Multisport fitness and wellness card with individual or family plans
- LuxMed healthcare coverage with individual or family plans
- UNUM life insurance with individual or family plans
- Onboarding allowance for necessary work equipment and setup
- Six self-care days in addition to standard Polish vacation entitlements
- Wellness and learning and development budgets
- Potential eligibility to purchase company stock or receive annual bonuses
Tech Stack
About PandaDoc
Stand out with the top‑rated solution for creating, managing, tracking, and esigning every important document you handle.
