QAD, Inc.

Sr. Site Reliability Engineer - SRE

QAD, Inc.
Apply
1 month ago
Remote, SpainSenior

Responsibilities

  • Design, implement, and maintain highly available, scalable, and resilient systems.
  • Build automation, self-service tooling, and software to reduce operational toil and improve reliability.
  • Own observability practices including monitoring, alerting, logging, tracing, dashboards, and synthetic testing.
  • Lead incident response, participate in on-call rotations, and drive blameless post-mortems.
  • Define, implement, and track SLIs, SLOs, and error budgets.
  • Use Terraform, Flux, and GitHub Actions for infrastructure as code, GitOps, and CI/CD automation.
  • Provide reliability expertise during system design reviews and influence architectural decisions.
  • Document processes, create runbooks, mentor engineers, and develop AI-assisted operational workflows with appropriate human oversight.

Requirements

  • Demonstrated experience operating and improving production systems at scale in an SRE, Production Engineering, or Platform Engineering role.
  • Ability to build mental models of complex distributed systems across infrastructure, applications, networking, identity, and observability.
  • Strong troubleshooting, incident response, and root cause analysis skills.
  • Experience defining and using SLIs, SLOs, and error budgets.
  • Experience with Kubernetes platforms including Amazon EKS, service meshes such as Istio, and AWS infrastructure and services.
  • Experience with identity and access management systems including Auth0 and AWS IAM.
  • Experience with GitOps workflows and infrastructure automation using Flux and Terraform.
  • Experience with observability platforms and practices, CI/CD systems, application logging, and distributed-system debugging.
  • Ability to build and maintain automation and tooling using one or more of Python, Go, or Bash.
  • Ability to lead complex incident follow-up, communicate clearly, and apply systems thinking to reliability improvements.
  • Ability to use and validate AI-assisted troubleshooting, root cause analysis, documentation, and operational workflows.
  • Experience with feature-flagging platforms such as LaunchDarkly is a bonus.

Benefits

  • Fully remote work arrangement.
  • Annual compensation range of 70,000–115,000 EUR.
  • Opportunity to shape SRE practices and improve operational excellence within a cloud-based enterprise software company.

Tech Stack

AWSBashDatadogGitHub ActionsGoIstioKubernetesPythonTerraform

Categories

Site Reliability
QAD, Inc.

About QAD, Inc.

1,001-5,000 employees
Contact me