
Senior Site Reliability Engineer (R-19383)
Dun & Bradstreet3 months ago
Dublin, IrelandSenior
Responsibilities
- Own and improve the reliability, availability, and performance of production services in Google Cloud Platform.
- Participate in incident management, including detection, triage, mitigation, escalation, and recovery.
- Improve incident workflows and tooling, including ServiceNow, to support clear ownership and timely communication.
- Design, implement, and operate observability solutions covering monitoring, logging, tracing, synthetics, and dashboards.
- Reduce operational toil through automation and drive SRE best practices.
- Support on-call rotations across multiple time zones as part of a 24/7 support model.
- Define, monitor, and report SLIs, SLOs, and error budgets for critical services.
- Drive service availability through automation, SRE principles, and proactive reliability engineering.
Requirements
- Bachelor’s degree in Computer Science, Information Technology, or a related field.
- Strong experience with cloud-native technologies, preferably Google Cloud Platform and Kubernetes/GKE.
- Proven experience with Site Reliability Engineering and production incident management, ideally using ServiceNow.
- Experience with monitoring and observability tools for metrics, logs, traces, and synthetics, including Splunk Observability and OpenTelemetry.
- Exposure to reliability testing, resilience engineering, or cost optimization initiatives.
- Software development or automation experience using Python, shell scripts, or similar languages.
- Hands-on experience operating production cloud infrastructure at scale.
- Experience managing multi-region, high-availability production systems with an emphasis on scalability, resilience, and minimizing service disruption.
- Proficiency with Microsoft Office Suites.
- Strong analytical, problem-solving, ownership, collaboration, curiosity, and proactive communication skills.
Tech Stack
Categories
Site Reliability