about 3 hours ago
Amsterdam, NetherlandsMid Level / Senior
Responsibilities
- Stand up and operate LLM and agent monitoring with Langfuse.
- Capture traces, latency, token usage, cost, quality scores, and safety signals.
- Build lightweight internal tooling and exporters in Python.
- Design and maintain Grafana dashboards and Prometheus metrics.
- Instrument platform and AI workloads for health, usage, cost, and SLA reporting.
- Feed telemetry and operational insights into the Platform Engineering backlog.
- Own Terraform IaC and CI/CD for observability tooling.
- Support incident investigation and root-cause analysis.
Requirements
- 5–8 years of experience in observability, SRE, platform engineering, DevOps, or cloud engineering.
- Strong experience with Azure Monitor, Application Insights, Log Analytics, and Managed Grafana.
- Hands-on experience with Langfuse, Grafana, and Prometheus.
- Experience with Terraform and CI/CD.
- Python skills for instrumentation, exporters, and automation.
- Familiarity with ML workloads and AI-specific metrics.
- Knowledge of logs, metrics, traces, dashboards, alerting, SLIs, and SLOs.
- Intermediate or higher English.
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment and talented teams