
Site Reliability Engineer- AI Enablement
Health Catalyst, Inc.2 months ago
Remote, United StatesSenior
Responsibilities
- Train and coach engineering teams on AI-assisted coding tools, prompt engineering, and agentic development patterns.
- Review AI system architectures for integration patterns, reliability risks, observability gaps, governance alignment, and operational readiness.
- Advise teams on AI governance, including model access, data handling, risk tiers, responsible AI use, prompt safety, and access controls.
- Provide hands-on solutioning and implementation guidance for LLM integrations, RAG pipelines, agentic architectures, and AI service patterns.
- Advise on observability, service-level objectives, failure modes, and reliability practices for AI-powered services, including incident support.
- Develop internal AI standards, reference architectures, reusable patterns, documentation, and review guidance.
- Collaborate with product managers, data scientists, security, compliance, and engineering stakeholders on regulatory and clinical requirements.
- Stay current with LLM capabilities, agentic frameworks, AI safety research, and SRE practices for AI systems.
Requirements
- At least 5 years of experience in site reliability engineering, platform engineering, or a closely related role.
- At least 2 years of hands-on experience solutioning or implementing AI or LLM-based systems in production or near-production contexts.
- Production experience with LLM API integration, including technologies such as Azure AI Foundry or Anthropic Claude.
- Hands-on experience with at least one agentic or RAG framework, such as LangChain, LlamaIndex, or Semantic Kernel.
- Strong SRE or platform engineering background with knowledge of observability, reliability principles, and operational practices.
- Experience evaluating AI architectures for reliability, security, governance alignment, and operational readiness.
- Experience advising, coaching, reviewing, or training engineering teams on AI tooling and best practices.
- Cloud infrastructure experience with Azure or AWS, including managed AI or ML services.
- Familiarity with Docker, Kubernetes, and CI/CD pipelines.
- Strong written and verbal communication skills and ability to explain complex AI concepts to varied audiences.
- Bachelor's or master's degree in Computer Science, Information Systems, or a related technical field, or equivalent practical experience.
- Preferred experience includes healthcare IT, HL7v2, CDA, EMR, FHIR, healthcare compliance, AI evaluation or red-teaming, rules engines, Datadog, Grafana, OpenTelemetry, Agile/Scrum, and Databricks.
- Ability to read and reason about software code and participate credibly in architecture discussions; hands-on coding is not a primary responsibility.
Benefits
- Remote work arrangement with no travel.
- The position is currently not eligible for visa sponsorship.
- Equal opportunity employer committed to respecting and benefiting from diverse backgrounds and experiences.
Tech Stack
Categories
Site Reliability