Hard Rock Digital, LLC

Senior Site Reliability Engineer

Hard Rock Digital, LLC
Apply
3 months ago
Remote, WorldwideSenior

Responsibilities

  • Maintain the availability, reliability, scalability, and performance of high-traffic Java applications in distributed environments.
  • Troubleshoot production and non-production issues, optimize JVM performance, and support performance testing and monitoring.
  • Deploy and manage Grafana, Prometheus, Loki, Mimir, and Alloy for monitoring, logging, alerting, and telemetry collection.
  • Create dashboards, alerts, and log queries and develop observability strategies for application and infrastructure health.
  • Design, build, and operate AI agents and agentic workflows for alert triage, root-cause analysis, runbook execution, incident response, and incident summarization.
  • Integrate LLM agents with Kubernetes, Grafana, Jira, Slack, and PagerDuty through tool calling and MCP servers.
  • Evaluate and operationalize agentic frameworks and implement guardrails, evaluation harnesses, and feedback loops.
  • Support incident response, post-mortems, root-cause analysis, and continuous improvement.
  • Build self-service tools and chatbot interfaces for querying system status, retrieving logs, and executing standard operating procedures.
  • Collaborate with developers, architects, data and ML engineers, DevOps teams, and NOC teams while mentoring junior team members.

Requirements

  • Bachelor's degree in Computer Science or a related field, or equivalent professional experience.
  • At least 5 years of experience in SRE, DevOps, or similar infrastructure roles managing large-scale, high-availability production systems.
  • At least 3 years of hands-on experience managing production Kubernetes clusters, including architecture, networking, storage, security, autoscaling, upgrades, and multi-cluster management.
  • Proficiency with kubectl, Helm, Kubernetes operators, container troubleshooting, Grafana, PromQL, Loki, and Grafana Alloy.
  • Hands-on experience managing distributed Java applications, including JVM tuning and optimization.
  • Expertise with AWS preferred, with GCP or Azure also valued.
  • Familiarity with Terraform, Terragrunt, Ansible, ArgoCD, GitOps, CI/CD pipelines, and deployment automation.
  • Strong scripting ability in Python, Bash, or Go.
  • Experience with on-call rotations, incident response, and root-cause analysis.
  • At least 1 year of practical experience building or operating AI/LLM-powered tools, agents, or workflows in production or production-adjacent environments.
  • Experience with tool calling, RAG, multi-step reasoning, LLM APIs, agentic orchestration frameworks, prompt engineering, MCP servers, AI-assisted coding tools, and human-in-the-loop design.
  • Preferred experience includes vector databases, LLM evaluation frameworks, open-source AI/ML or SRE tooling, and data engineering or ML pipelines.
  • Strong communication, problem-solving, automation, collaboration, mentoring, and continuous-improvement skills.

Benefits

  • Competitive pay and benefits.
  • Flexible vacation allowance.
  • Hybrid or remote working environment.
  • Startup culture backed by a secure, global brand.
  • Inclusive and equal-opportunity work environment.

Tech Stack

Categories

AI ApplicationsSite Reliability
Hard Rock Digital, LLC

About Hard Rock Digital, LLC

1,001-5,000 employees
Contact me