
Site Reliability Engineer, AI & Agentic Systems
ServiceLink2 years ago
Plano, TX, USASenior
Responsibilities
- Own end-to-end reliability, availability, fault tolerance, and graceful degradation for large-scale Azure-hosted production systems.
- Lead production incident troubleshooting, root cause analysis, post-incident reviews, on-call response, and actionable reliability improvements.
- Define and enforce SLIs, SLOs, error budgets, failure-mode analysis, capacity planning, and operational standards.
- Build and operate resilient Azure services and observability platforms using Kubernetes, Prometheus, Loki, Tempo, and Grafana.
- Create operational automation, self-healing mechanisms, automated remediation workflows, and runbook automation to reduce toil and MTTR.
- Manage API lifecycle and traffic management with Gravitee API Gateway and build durable workflows using Temporal.
- Administer and tune PostgreSQL databases for reliability, performance, replication, backup, recovery, and high availability.
- Design and execute load, stress, soak, endurance, spike, and capacity testing for distributed systems and microservices.
- Develop Micro Focus LoadRunner and VuGen performance test scripts, analyze results, and recommend improvements for bottlenecks and scalability limits.
- Integrate performance testing into CI/CD pipelines and establish performance baselines, benchmarks, and SLAs.
- Design and implement AI-driven and agentic automation for incident triage, alert correlation, diagnosis, remediation, and predictive operational insights.
- Integrate AI agents with monitoring, CI/CD, and incident management systems while ensuring guardrails, fallback mechanisms, observability, and audit trails.
- Mentor junior SREs, participate in architecture reviews, and influence reliability, observability, performance, and non-functional engineering practices.
Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, or Production Engineering roles.
- Strong hands-on experience troubleshooting distributed systems in production at scale.
- Knowledge of Linux internals, TCP/IP, DNS, HTTP, TLS, and system performance tuning.
- Deep hands-on experience with Microsoft Azure, including compute, networking, storage, managed services, and AKS.
- Strong knowledge of Kubernetes, container orchestration, Helm charts, and microservices architectures.
- Proficiency in Python, Go, Java, or an equivalent programming language.
- Experience with Azure DevOps or GitHub Actions CI/CD pipelines and Terraform, ARM Templates, or Bicep infrastructure as code.
- Hands-on experience operating Prometheus, Grafana, Loki, and Tempo observability stacks.
- Experience with alerting strategies, SLI/SLO monitoring, and on-call incident management.
- Proven experience designing and executing performance and load testing for large-scale distributed applications.
- Hands-on proficiency with Micro Focus LoadRunner and VuGen, including virtual-user scripting, parameterization, correlation, and result analysis.
- Understanding of load, stress, endurance/soak, spike, and capacity-planning methodologies and performance metrics.
- Experience integrating performance tests into automated CI/CD pipelines.
- Experience with Gravitee or an equivalent API gateway platform.
- Hands-on experience with Temporal for workflow orchestration and durable execution.
- Strong PostgreSQL administration skills, including query optimization, replication, backup/recovery, and performance tuning.
- Hands-on experience building or integrating AI-powered automation in production environments.
- Experience with agent-based systems, LLM-powered workflows, Retrieval-Augmented Generation, or intelligent assistants.
- Familiarity with Azure OpenAI, Cognitive Services, and Azure ML, plus understanding of reliability and safety challenges for AI systems in production.
Benefits
- Hybrid work arrangement based at the Plano, Texas office, with in-office work required three days per week.
- Applicants must be currently authorized to work full-time in the United States and must not require employment visa sponsorship now or in the future.
Tech Stack
Categories
AI ApplicationsSite Reliability