
Senior Platform Engineer — AI Infrastructure
Helloprint3 days ago
Rotterdam, NetherlandsSenior
Responsibilities
- Define and enforce SLOs, SLIs, and error-budget policies across critical customer journeys and services.
- Expand distributed observability, tracing, telemetry, and automated diagnostics across microservices, queue workers, and cloud infrastructure.
- Develop progressive delivery workflows with canary traffic shifting, automated health gates, and SLO-driven rollbacks.
- Architect and operate runtime infrastructure for AI agents, semantic pipelines, and background automation.
- Manage AI runtime costs, model and token budgets, latency, rate limits, queue backpressure, and provider availability.
- Drive FinOps practices and optimize Google Cloud compute, storage, and infrastructure budgets.
- Lead on-call incident response and blameless post-mortems, converting root causes into automated tests, synthetic checks, and guardrails.
- Own capacity forecasting, dependency isolation, load testing, disaster recovery validation, and RTO/RPO readiness.
- Manage Terraform-based infrastructure workflows, Google Cloud Run environments, IAM least privilege, and Secret Manager.
- Build internal platform tooling, runbooks, and self-service deployment capabilities to reduce toil and improve developer experience.
Requirements
- Proven experience operating and scaling high-traffic distributed production systems with demanding uptime, latency, and release-safety requirements.
- Strong troubleshooting skills across Linux, containerized runtimes, relational and NoSQL databases, Redis queues, and cloud network boundaries.
- Extensive hands-on experience with Google Cloud Platform, Google Cloud Run, and Terraform.
- Demonstrated ability to monitor, analyze, and optimize cloud and AI runtime costs.
- Practical experience configuring Sentry and Google Cloud Monitoring for error reporting, tracing, SLI tracking, and canary gates.
- Proficiency in Python or TypeScript/JavaScript for platform tooling.
- Practical working familiarity with modern PHP in a Laravel ecosystem.
- Familiarity with operationalizing LLM integrations, embeddings and vector workflows, rate-limited external APIs, or background task orchestration.
- Pragmatic engineering approach emphasizing durable guardrails, deterministic automation, and reduced manual intervention.
Benefits
- Opportunities to grow within the role and across HelloPrint.
- International work environment with more than 100 professionals from 20+ nationalities in Rotterdam or Valencia.
- Opportunity to shape a high-growth, AI-driven company with substantial ownership and impact.
- Freedom to take initiative and drive projects with experienced leaders.
- 24/7 HelloFit gym access, UrbanSportsClub discount, company events, and company-sponsored healthy meals every day.
- Work location is Rotterdam or Valencia.
Tech Stack
Categories
DevOpsSite Reliability