4 months ago
Columbus, OH, USAStaff+

Responsibilities

  • Design, develop, and implement AI-driven systems and automation tools that improve the reliability and efficiency of digital platforms.
  • Monitor application and infrastructure health, availability, and performance using SRE practices and observability tools.
  • Integrate machine learning models into production environments and ensure reliable deployment and operation.
  • Establish and enforce SLOs, error budgets, monitoring, alerting, and incident response procedures for AI-driven services.
  • Troubleshoot complex AI-system incidents, conduct post-incident analysis, automate manual tasks, and optimize performance.
  • Develop abstraction layers across AI providers including Google and OpenAI.
  • Conduct design workshops, proofs of concept, and code-with sessions to shape data-driven agent workflows with stakeholders.
  • Define metrics, test harnesses, and evaluation plans for agent accuracy, latency, safety, and cost effectiveness.
  • Create reusable patterns, documentation, runbooks, and best practices for internal assets and client roadmaps.

Requirements

  • Bachelor’s or master’s degree in computer science, engineering, data science, or a related field.
  • Minimum five years of experience in AI/ML engineering, SRE, DevOps, or related roles.
  • Strong programming skills in Python, Java, or similar languages, including experience developing and deploying machine learning models.
  • Hands-on experience with AWS, GCP, or Azure and containerization technologies such as Docker and Kubernetes.
  • Familiarity with Prometheus, Grafana, the ELK stack, and ServiceNow incident management platforms.
  • Solid understanding of SRE principles, including monitoring, alerting, SLOs, error budgets, and automation.
  • Experience with infrastructure-as-code tools such as Terraform and Ansible and with CI/CD pipelines.
  • Preferred experience operationalizing large language models or generative AI systems in production.
  • Preferred background in MLOps, data engineering, and/or cloud-native AI deployment.
  • Strong communication, documentation, problem-solving, and collaboration abilities are required; security best practices and contributions to open-source AI/SRE communities are preferred.

Benefits

  • The position is office-based, with certain arrangements potentially eligible for a flexible combination of in-office and work-from-home work; specific arrangements are provided by the hiring team.
  • The role is exempt and may involve on-call responsibilities.
  • Huntington is an Equal Opportunity Employer and maintains a tobacco-free hiring practice.

Categories

AI ApplicationsSite Reliability
Huntington Bancshares Incorporated

About Huntington Bancshares Incorporated

10,000+ employees
Contact me