
Senior Site Reliability Engineer
Hard Rock Digital, LLC3 months ago
Remote, WorldwideSenior
Responsibilities
- Maintain the availability, reliability, scalability, and performance of high-traffic Java applications in distributed environments.
- Troubleshoot production and non-production issues, optimize JVM performance, and support performance testing and monitoring.
- Deploy and manage Grafana, Prometheus, Loki, Mimir, and Alloy for monitoring, logging, alerting, and telemetry collection.
- Create dashboards, alerts, and log queries and develop observability strategies for application and infrastructure health.
- Design, build, and operate AI agents and agentic workflows for alert triage, root-cause analysis, runbook execution, incident response, and incident summarization.
- Integrate LLM agents with Kubernetes, Grafana, Jira, Slack, and PagerDuty through tool calling and MCP servers.
- Evaluate and operationalize agentic frameworks and implement guardrails, evaluation harnesses, and feedback loops.
- Support incident response, post-mortems, root-cause analysis, and continuous improvement.
- Build self-service tools and chatbot interfaces for querying system status, retrieving logs, and executing standard operating procedures.
- Collaborate with developers, architects, data and ML engineers, DevOps teams, and NOC teams while mentoring junior team members.
Requirements
- Bachelor's degree in Computer Science or a related field, or equivalent professional experience.
- At least 5 years of experience in SRE, DevOps, or similar infrastructure roles managing large-scale, high-availability production systems.
- At least 3 years of hands-on experience managing production Kubernetes clusters, including architecture, networking, storage, security, autoscaling, upgrades, and multi-cluster management.
- Proficiency with kubectl, Helm, Kubernetes operators, container troubleshooting, Grafana, PromQL, Loki, and Grafana Alloy.
- Hands-on experience managing distributed Java applications, including JVM tuning and optimization.
- Expertise with AWS preferred, with GCP or Azure also valued.
- Familiarity with Terraform, Terragrunt, Ansible, ArgoCD, GitOps, CI/CD pipelines, and deployment automation.
- Strong scripting ability in Python, Bash, or Go.
- Experience with on-call rotations, incident response, and root-cause analysis.
- At least 1 year of practical experience building or operating AI/LLM-powered tools, agents, or workflows in production or production-adjacent environments.
- Experience with tool calling, RAG, multi-step reasoning, LLM APIs, agentic orchestration frameworks, prompt engineering, MCP servers, AI-assisted coding tools, and human-in-the-loop design.
- Preferred experience includes vector databases, LLM evaluation frameworks, open-source AI/ML or SRE tooling, and data engineering or ML pipelines.
- Strong communication, problem-solving, automation, collaboration, mentoring, and continuous-improvement skills.
Benefits
- Competitive pay and benefits.
- Flexible vacation allowance.
- Hybrid or remote working environment.
- Startup culture backed by a secure, global brand.
- Inclusive and equal-opportunity work environment.
Tech Stack
Categories
AI ApplicationsSite Reliability