3 months ago
Remote, CanadaSenior
Responsibilities
- Own end-to-end development of multi-agent AI systems, including architecture, implementation, testing, deployment, and ongoing operation.
- Build modular agentic systems and reusable agentic skills accessible through Slack, dashboards, internal applications, and CLIs.
- Implement observability and feedback loops covering logging, performance metrics, prompt iteration, model evaluation, and cost management.
- Establish governance and compliance standards for AI workflows, including access controls, audit trails, PII handling, and human escalation.
- Build MCP servers, APIs, CLIs, and microservices that connect AI models to BigQuery, Slack, CRMs, email, calendars, and analytics tools.
- Architect RAG data flows connecting LLMs to internal knowledge bases, customer data, and real-time business context.
- Build scalable serverless or containerized services using Cloud Functions and Cloud Run.
- Partner with RevOps, Demand Generation, Regional Marketing, and SDR teams to identify and solve high-impact automation problems.
- Design and deploy reliable workflows using n8n, Workato, or custom platforms.
- Create documentation, playbooks, and enablement materials so partner teams can operate automation independently.
- Define the automation platform’s technical direction, including data models, API contracts, shared libraries, and reference architectures.
Requirements
- 8+ years of software engineering experience with depth in backend development, systems integration, or data and analytics engineering.
- 2+ years of hands-on experience applying LLMs or AI to production workflows.
- Strong proficiency in Python and JavaScript/Node.js, with Git-based workflows, code review, and testing experience.
- Experience with LLM frameworks and patterns including prompt engineering, retrieval-augmented generation, function calling or tool use, structured output parsing, and evaluation.
- Experience building and operating multi-agent systems at scale, including orchestration, state management, and production monitoring.
- Deep familiarity with Google Cloud Platform, BigQuery, Cloud Functions, and Cloud Run.
- Understanding of LLM failure modes and mitigations such as confidence thresholds, fallback logic, human escalation, and cost or latency management.
- Ability to independently identify high-impact business problems, challenge low-impact requests, and deliver end-to-end solutions.
- Fluency with AI-assisted development tools such as GitHub Copilot, Cursor, and Claude Code.
- Clear technical communication skills for explaining complex systems to technical and business stakeholders.
- Bonus experience may include vector databases or retrieval pipelines, marketing or sales platforms, frontend interfaces, AI observability tooling, workflow orchestration platforms, Model Context Protocol, B2B SaaS workflow automation, and open-source communities.
Benefits
- 100% remote work for candidates in Canada, excluding Quebec residents.
- In-person onboarding is provided.
- 30 days of annual leave per year, including 3 Grafana Shutdown Days, subject to local legislation.
- Restricted Stock Units (RSUs) are included.
- Career growth pathways and a global, collaborative culture are offered.
Tech Stack
Categories
About Grafana
Grafana Labs builds open-source observability tools and a managed SaaS, Grafana Cloud, used by engineering and SRE teams to monitor, visualize, and analyze telemetry. Its products include Grafana dashboards plus Loki (logs), Tempo (traces), and Mimir (metrics), offered as cloud subscriptions and enterprise software. Privately held and headquartered in New York, it serves customers such as Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce.
