Hard Rock Digital, LLC

Senior Site Reliability Engineer

Hard Rock Digital, LLC
Apply
3 months ago
Remote, PolandSenior

Responsibilities

  • Maintain the availability, reliability, scalability, and performance of high-traffic Java applications in distributed production environments.
  • Troubleshoot application and infrastructure issues, perform JVM tuning, and support performance testing and optimization.
  • Deploy and manage Grafana, Prometheus, Loki, Mimir, and Alloy for monitoring, logging, alerting, and telemetry.
  • Create dashboards, alerts, log queries, and observability strategies for application and infrastructure health.
  • Design and operate AI agents and agentic workflows for alert triage, root-cause analysis, runbook execution, incident response, and incident summarization.
  • Integrate LLM agents with Kubernetes, Grafana, Jira, Slack, PagerDuty, and other infrastructure APIs using tool calling and human approval workflows.
  • Build and maintain MCP servers and integrations that expose internal systems to AI agents.
  • Evaluate and operationalize agentic frameworks and workflow platforms, and implement guardrails, evaluation harnesses, and feedback loops.
  • Participate in incident response, post-mortems, root-cause analysis, and continuous improvement initiatives.
  • Build self-service tools and chatbot interfaces for querying system status, retrieving logs, and executing standard operating procedures.
  • Measure operational toil reduction and collaborate with developers, architects, data and ML engineers, DevOps teams, and NOC teams.

Requirements

  • A degree in Computer Science or a related field, or equivalent professional experience.
  • At least 5 years of experience in SRE, DevOps, or similar infrastructure roles managing large-scale, high-availability production systems.
  • At least 3 years of hands-on experience managing production Kubernetes clusters, including architecture, networking, storage, security, autoscaling, upgrades, and multi-cluster management.
  • Proficiency with kubectl, Helm, Kubernetes operators, container troubleshooting, Grafana, PromQL, Loki, and Grafana Alloy.
  • Hands-on experience managing Java applications in distributed environments, including JVM tuning and optimization.
  • Expertise with AWS preferred, with GCP or Azure also valued.
  • Familiarity with Terraform, Terragrunt, Ansible, and ArgoCD, plus scripting in Python, Bash, or Go and experience with CI/CD and deployment automation.
  • Experience with on-call rotations, incident response, and root-cause analysis.
  • At least 1 year of practical experience building or operating AI/LLM-powered tools, agents, or workflows in production or production-adjacent environments.
  • Experience with tool calling, RAG, multi-step reasoning, LLM APIs, agentic orchestration frameworks, prompt engineering, MCP servers, and AI-assisted coding tools.
  • Understanding of AI safety, hallucination mitigation, human-in-the-loop design, and operational safeguards.
  • Preferred qualifications include vector database experience, LLM evaluation frameworks, open-source AI/ML or SRE contributions, and data engineering or ML pipeline experience.
  • Strong communication, problem-solving, mentoring, automation, and continuous-improvement abilities.

Benefits

  • Fully remote work from Poland only.
  • Full-time B2B engagement.
  • Autonomy to experiment with AI tools and leadership support for deploying them in production.

Tech Stack

Categories

AI ApplicationsSite Reliability
Hard Rock Digital, LLC

About Hard Rock Digital, LLC

1,001-5,000 employees
Contact me