
Senior Site Reliability Engineer (SRE & AI Platform Operations)
Helloprint7 days ago
Rotterdam, NetherlandsSenior
Responsibilities
- Define, track, and enforce SLOs, SLIs, and error-budget policies for critical customer journeys and services.
- Expand distributed observability, tracing, telemetry, and automated diagnostics across microservices, queue workers, and Google Cloud infrastructure.
- Build deployment safety practices including canary traffic shifting, automated health gates, and SLO-driven rollbacks in GitHub Actions.
- Operate and scale AI runtime infrastructure while managing model and token budgets, latency, rate limits, queue backpressure, and provider availability.
- Drive FinOps practices by optimizing cloud resources, infrastructure budgets, and AI runtime costs.
- Lead on-call incident response and blameless post-mortems, converting root causes into automated tests, synthetic checks, and architectural guardrails.
- Drive capacity forecasting, dependency isolation, load testing, and disaster recovery validation against RTO/RPO targets.
- Own Terraform and Google Cloud Run infrastructure workflows, IAM least privilege, Secret Manager usage, and deterministic environments.
- Build internal tooling, runbooks, and self-service deployment primitives that reduce toil and improve developer experience.
Requirements
- Proven experience operating and scaling high-traffic distributed production systems with demanding uptime, latency, and release-safety requirements.
- Extensive hands-on experience with Google Cloud Platform, Google Cloud Run, and Terraform-based infrastructure automation.
- Strong troubleshooting skills across Linux environments, containerized runtimes, relational and NoSQL databases, Redis queues, and cloud network boundaries.
- Demonstrated ability to monitor, analyze, and optimize cloud and AI runtime costs while balancing performance, reliability, and budget efficiency.
- Practical experience configuring Sentry and Google Cloud Monitoring for error reporting, distributed tracing, SLI tracking, and canary gates.
- Proficiency in Python or TypeScript/JavaScript for platform tooling and practical familiarity with modern PHP in a Laravel ecosystem.
- Familiarity with operationalizing LLM integrations, embeddings/vector workflows, rate-limited external APIs, or background task orchestration.
- Pragmatic engineering mindset with strong technical agency and a preference for durable guardrails and deterministic automation.
Benefits
- Opportunities to grow within the role and across HelloPrint.
- International work environment with more than 100 professionals from 20+ nationalities in Rotterdam or Valencia.
- Opportunity to shape a high-growth, AI-driven company with visible impact.
- Ownership from day one, experienced leadership, and freedom to innovate.
- HelloBenefits including 24/7 HelloFit gym access, an UrbanSportsClub discount, company events, and company-sponsored healthy meals every day.
Tech Stack
Categories
DevOpsSite Reliability