ServiceLink

Site Reliability Engineer, AI & Agentic Systems

ServiceLink
Apply
2 years ago
Plano, TX, USASenior

Responsibilities

  • Own end-to-end reliability, availability, fault tolerance, and graceful degradation for large-scale Azure-hosted production systems.
  • Lead production incident troubleshooting, root cause analysis, post-incident reviews, on-call response, and actionable reliability improvements.
  • Define and enforce SLIs, SLOs, error budgets, failure-mode analysis, capacity planning, and operational standards.
  • Build and operate resilient Azure services and observability platforms using Kubernetes, Prometheus, Loki, Tempo, and Grafana.
  • Create operational automation, self-healing mechanisms, automated remediation workflows, and runbook automation to reduce toil and MTTR.
  • Manage API lifecycle and traffic management with Gravitee API Gateway and build durable workflows using Temporal.
  • Administer and tune PostgreSQL databases for reliability, performance, replication, backup, recovery, and high availability.
  • Design and execute load, stress, soak, endurance, spike, and capacity testing for distributed systems and microservices.
  • Develop Micro Focus LoadRunner and VuGen performance test scripts, analyze results, and recommend improvements for bottlenecks and scalability limits.
  • Integrate performance testing into CI/CD pipelines and establish performance baselines, benchmarks, and SLAs.
  • Design and implement AI-driven and agentic automation for incident triage, alert correlation, diagnosis, remediation, and predictive operational insights.
  • Integrate AI agents with monitoring, CI/CD, and incident management systems while ensuring guardrails, fallback mechanisms, observability, and audit trails.
  • Mentor junior SREs, participate in architecture reviews, and influence reliability, observability, performance, and non-functional engineering practices.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering roles.
  • Strong hands-on experience troubleshooting distributed systems in production at scale.
  • Knowledge of Linux internals, TCP/IP, DNS, HTTP, TLS, and system performance tuning.
  • Deep hands-on experience with Microsoft Azure, including compute, networking, storage, managed services, and AKS.
  • Strong knowledge of Kubernetes, container orchestration, Helm charts, and microservices architectures.
  • Proficiency in Python, Go, Java, or an equivalent programming language.
  • Experience with Azure DevOps or GitHub Actions CI/CD pipelines and Terraform, ARM Templates, or Bicep infrastructure as code.
  • Hands-on experience operating Prometheus, Grafana, Loki, and Tempo observability stacks.
  • Experience with alerting strategies, SLI/SLO monitoring, and on-call incident management.
  • Proven experience designing and executing performance and load testing for large-scale distributed applications.
  • Hands-on proficiency with Micro Focus LoadRunner and VuGen, including virtual-user scripting, parameterization, correlation, and result analysis.
  • Understanding of load, stress, endurance/soak, spike, and capacity-planning methodologies and performance metrics.
  • Experience integrating performance tests into automated CI/CD pipelines.
  • Experience with Gravitee or an equivalent API gateway platform.
  • Hands-on experience with Temporal for workflow orchestration and durable execution.
  • Strong PostgreSQL administration skills, including query optimization, replication, backup/recovery, and performance tuning.
  • Hands-on experience building or integrating AI-powered automation in production environments.
  • Experience with agent-based systems, LLM-powered workflows, Retrieval-Augmented Generation, or intelligent assistants.
  • Familiarity with Azure OpenAI, Cognitive Services, and Azure ML, plus understanding of reliability and safety challenges for AI systems in production.

Benefits

  • Hybrid work arrangement based at the Plano, Texas office, with in-office work required three days per week.
  • Applicants must be currently authorized to work full-time in the United States and must not require employment visa sponsorship now or in the future.

Tech Stack

AzureGitHub ActionsGoGrafanaHelmJavaKubernetesPostgreSQLPrometheusPythonTerraform

Categories

AI ApplicationsSite Reliability
ServiceLink

About ServiceLink

1,001-5,000 employees
Contact me