Helloprint

Senior Site Reliability Engineer (SRE & AI Platform Operations)

Helloprint
Apply
7 days ago
Rotterdam, NetherlandsSenior

Responsibilities

  • Define, track, and enforce SLOs, SLIs, and error-budget policies for critical customer journeys and services.
  • Expand distributed observability, tracing, telemetry, and automated diagnostics across microservices, queue workers, and Google Cloud infrastructure.
  • Build deployment safety practices including canary traffic shifting, automated health gates, and SLO-driven rollbacks in GitHub Actions.
  • Operate and scale AI runtime infrastructure while managing model and token budgets, latency, rate limits, queue backpressure, and provider availability.
  • Drive FinOps practices by optimizing cloud resources, infrastructure budgets, and AI runtime costs.
  • Lead on-call incident response and blameless post-mortems, converting root causes into automated tests, synthetic checks, and architectural guardrails.
  • Drive capacity forecasting, dependency isolation, load testing, and disaster recovery validation against RTO/RPO targets.
  • Own Terraform and Google Cloud Run infrastructure workflows, IAM least privilege, Secret Manager usage, and deterministic environments.
  • Build internal tooling, runbooks, and self-service deployment primitives that reduce toil and improve developer experience.

Requirements

  • Proven experience operating and scaling high-traffic distributed production systems with demanding uptime, latency, and release-safety requirements.
  • Extensive hands-on experience with Google Cloud Platform, Google Cloud Run, and Terraform-based infrastructure automation.
  • Strong troubleshooting skills across Linux environments, containerized runtimes, relational and NoSQL databases, Redis queues, and cloud network boundaries.
  • Demonstrated ability to monitor, analyze, and optimize cloud and AI runtime costs while balancing performance, reliability, and budget efficiency.
  • Practical experience configuring Sentry and Google Cloud Monitoring for error reporting, distributed tracing, SLI tracking, and canary gates.
  • Proficiency in Python or TypeScript/JavaScript for platform tooling and practical familiarity with modern PHP in a Laravel ecosystem.
  • Familiarity with operationalizing LLM integrations, embeddings/vector workflows, rate-limited external APIs, or background task orchestration.
  • Pragmatic engineering mindset with strong technical agency and a preference for durable guardrails and deterministic automation.

Benefits

  • Opportunities to grow within the role and across HelloPrint.
  • International work environment with more than 100 professionals from 20+ nationalities in Rotterdam or Valencia.
  • Opportunity to shape a high-growth, AI-driven company with visible impact.
  • Ownership from day one, experienced leadership, and freedom to innovate.
  • HelloBenefits including 24/7 HelloFit gym access, an UrbanSportsClub discount, company events, and company-sponsored healthy meals every day.

Tech Stack

AstroGitHub ActionsGoogle CloudJavaScriptLaravelLinuxPHPPythonRedisTerraformTypeScript

Categories

DevOpsSite Reliability
Helloprint

About Helloprint

51-200 employees
Contact me