Charles Schwab

Senior AI Site Reliability Engineer, AI.x

Charles Schwab
Apply
7 days ago

Responsibilities

  • Lead automation-first reliability initiatives and define roadmaps for observability and self-healing systems across AI.x platforms.
  • Design and implement automated CI/CD pipelines with testing, validation, rollback, and one-touch deployment capabilities.
  • Build observability frameworks for AI services using metrics, logs, traces, intelligent alerting, and automated diagnostics.
  • Participate in a 24/7 on-call rotation, leading incident response, root cause analysis, and production issue resolution.
  • Establish and manage SLOs, SLIs, error budgets, and incident response runbooks.
  • Automate infrastructure provisioning, configuration management, and deployment through Infrastructure as Code practices.
  • Partner with AI Engineering teams to integrate reliability practices throughout the development lifecycle.
  • Identify and resolve reliability, performance, and scalability issues through analysis, capacity planning, and system optimization.
  • Maintain monitoring, alerting, and incident response frameworks for production AI systems and data pipelines.

Requirements

  • 8+ years of software engineering experience, including 4+ years as a hands-on Site Reliability Engineer.
  • Bachelor’s degree in Computer Science or a related field, or equivalent experience.
  • 5+ years of building complex products from scratch, operating them in production, and ensuring their reliability.
  • 3+ years working with containers and cloud-native applications in public-cloud environments using Infrastructure as Code and CI/CD pipelines.
  • 3+ years working in high-availability hybrid-cloud environments.
  • Strong computer science fundamentals and experience across the technology stack.
  • Experience with proprietary or open-source LLMs such as Gemini, Claude, or OpenAI, including deploying LLM-powered applications to production.
  • Strong understanding of observability, incident management, and reliability engineering principles.
  • Ability to troubleshoot complex distributed-system problems with ambiguous or incomplete data.
  • Experience with Terraform and Google Cloud Platform is preferred.

Benefits

  • The role is eligible for bonus or incentive opportunities in addition to the salary range.
  • Participation in a 24/7 on-call rotation is required.
  • The position is part of Schwab’s AI Strategy & Transformation team based in San Francisco.

Categories

Site Reliability
Charles Schwab

About Charles Schwab

10,000+ employees

Charles Schwab is a different kind of investment services firm – one that strives to disrupt the status quo of the traditional Wall Street approach on behalf of our clients. We believe today, as we did on Day 1, that when you find ways to improve the investing experience for your clients, then business results will follow. Follow our company culture at #SchwabLife and see how we give back at #Schwab4Good. Support hours: 7 a.m.–7 p.m. CT or 24/7 at schwab.com/contact-us. Social Media Disclosures: https://www.aboutschwab.com/social-media (#0424-TM8W)

Contact me