
Platform Reliability Engineer
Apify Technologies s.r.o.2 months ago
Prague, CzechiaSenior
Responsibilities
- Operate and improve monitoring using Prometheus, Grafana, and OpenTelemetry.
- Define production metrics, dashboards, monitoring standards, and actionable alerting with appropriate severity, routing, and noise reduction.
- Support incident processes from detection through resolution, communication, documentation, post-incident review, and follow-up improvements.
- Create and maintain status pages, runbooks, incident documentation, observability guidance, and alerting principles.
- Partner with platform and product engineers to adopt reliability tooling and practices.
- Automate repetitive tasks and improve developer workflows and infrastructure support.
- Own the monitoring and alerting improvement roadmap and align priorities with leadership.
- Participate in team ceremonies, technical discussions, incident reviews, and infrastructure initiatives.
Requirements
- Hands-on experience selecting production signals that reflect customer experience.
- Experience working with incidents and alerts from detection through resolution and follow-up.
- Hands-on experience with Prometheus, Grafana, OpenTelemetry, or similar tools and alert-routing tools such as PagerDuty.
- Ability to read and write code, follow services and pipelines across the stack, and collaborate on technical implementation details.
- Practical understanding of blame-free, learning-focused post-incident culture.
- Ability to write clear, concise guidance that engineering teams adopt.
- Motivation to automate repetitive tasks and improve developer workflows.
- Preferred: meaningful hands-on application or backend development experience with production systems.
- Preferred: experience building and maintaining AWS infrastructure, including EC2, EKS, S3, or CloudFormation, and experience with container technologies.
- Preferred: familiarity with CI/CD pipelines or release practices.
Benefits
- Full-time position based in Prague or Brno, with an option to work remotely.
- Flexible working hours and no fixed holiday-counting policy as long as work is completed.
- Stock options and profit sharing.
- Personal growth support, education and training budget, conference tickets, and cross-team work opportunities.
- Unlimited Claude access for employees.
- Generous hardware budget.
- Free office lunches, drinks, snacks, and access to office recreational activities.
- Epic team buildings and offsites, free Prague Zoo entry, and a free Multisport card.
- Welcoming offices for pets, children, and bikes.
Tech Stack
Amazon DynamoDBAWSCypressExpressGitHub ActionsGrafanaHelmKubernetesMongoDBNestJSNext.jsNode.jsPrometheusReactRedisTypeScript
Categories
DevOpsSite Reliability