
Senior Site Reliability Engineer
Hard Rock Digital, LLC3 months ago
Remote, PolandSenior
Responsibilities
- Maintain the availability, reliability, scalability, and performance of high-traffic Java applications in distributed production environments.
- Troubleshoot application and infrastructure issues, perform JVM tuning, and support performance testing and optimization.
- Deploy and manage Grafana, Prometheus, Loki, Mimir, and Alloy for monitoring, logging, alerting, and telemetry.
- Create dashboards, alerts, log queries, and observability strategies for application and infrastructure health.
- Design and operate AI agents and agentic workflows for alert triage, root-cause analysis, runbook execution, incident response, and incident summarization.
- Integrate LLM agents with Kubernetes, Grafana, Jira, Slack, PagerDuty, and other infrastructure APIs using tool calling and human approval workflows.
- Build and maintain MCP servers and integrations that expose internal systems to AI agents.
- Evaluate and operationalize agentic frameworks and workflow platforms, and implement guardrails, evaluation harnesses, and feedback loops.
- Participate in incident response, post-mortems, root-cause analysis, and continuous improvement initiatives.
- Build self-service tools and chatbot interfaces for querying system status, retrieving logs, and executing standard operating procedures.
- Measure operational toil reduction and collaborate with developers, architects, data and ML engineers, DevOps teams, and NOC teams.
Requirements
- A degree in Computer Science or a related field, or equivalent professional experience.
- At least 5 years of experience in SRE, DevOps, or similar infrastructure roles managing large-scale, high-availability production systems.
- At least 3 years of hands-on experience managing production Kubernetes clusters, including architecture, networking, storage, security, autoscaling, upgrades, and multi-cluster management.
- Proficiency with kubectl, Helm, Kubernetes operators, container troubleshooting, Grafana, PromQL, Loki, and Grafana Alloy.
- Hands-on experience managing Java applications in distributed environments, including JVM tuning and optimization.
- Expertise with AWS preferred, with GCP or Azure also valued.
- Familiarity with Terraform, Terragrunt, Ansible, and ArgoCD, plus scripting in Python, Bash, or Go and experience with CI/CD and deployment automation.
- Experience with on-call rotations, incident response, and root-cause analysis.
- At least 1 year of practical experience building or operating AI/LLM-powered tools, agents, or workflows in production or production-adjacent environments.
- Experience with tool calling, RAG, multi-step reasoning, LLM APIs, agentic orchestration frameworks, prompt engineering, MCP servers, and AI-assisted coding tools.
- Understanding of AI safety, hallucination mitigation, human-in-the-loop design, and operational safeguards.
- Preferred qualifications include vector database experience, LLM evaluation frameworks, open-source AI/ML or SRE contributions, and data engineering or ML pipeline experience.
- Strong communication, problem-solving, mentoring, automation, and continuous-improvement abilities.
Benefits
- Fully remote work from Poland only.
- Full-time B2B engagement.
- Autonomy to experiment with AI tools and leadership support for deploying them in production.
Tech Stack
Categories
AI ApplicationsSite Reliability