Capgemini

AI Platform & Site Reliability Engineering Managing Consultant

Capgemini
Apply
6 days ago
London, United KingdomStaff+
H1B Sponsor

Responsibilities

  • Develop enterprise AI platform strategies and architectures covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations, and integration patterns.
  • Design and implement model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation, and operational guardrails.
  • Establish SRE practices including SLIs, SLOs, error budgets, capacity planning, resilience engineering, incident management, and reliability governance.
  • Define observability strategies across applications, platforms, and AI workloads using metrics, logs, traces, and telemetry.
  • Monitor AI-specific failure modes including data-quality degradation, hallucination patterns, token consumption, agent reliability, and model-performance drift.
  • Design AI governance, model-risk controls, compliance measures, and responsible AI practices addressing security, regulatory, and ethical requirements.
  • Lead client assessments, transformation roadmaps, implementation workstreams, multidisciplinary delivery teams, propositions, bids, and business-growth activities.

Requirements

  • Proven experience designing, delivering, and operating cloud-native, platform engineering, AI platform, or reliability engineering solutions in complex enterprise environments.
  • Strong understanding of AI platform architectures, LLMOps, MLOps, agentic AI frameworks, model lifecycle management, and AI operational controls.
  • Experience establishing and scaling SRE practices including observability, SLIs, SLOs, error budgets, incident management, and reliability engineering.
  • Strong understanding of AI governance, responsible AI, regulatory requirements, and model risk management.
  • Experience implementing observability strategies with modern monitoring, telemetry, and operational analytics platforms.
  • Demonstrated ability to advise senior stakeholders and lead multidisciplinary transformation programmes.
  • Experience across Azure, AWS, and Google Cloud Platform ecosystems.
  • Ability to balance business outcomes, user needs, engineering constraints, and operational requirements when shaping platform strategies.
  • Desirable experience with AI observability, model monitoring, AI governance tooling, platform engineering, DevSecOps, automation, and Infrastructure-as-Code.
  • Desirable certifications include Azure AI Engineer Associate, Azure Solutions Architect Expert, AWS Machine Learning Specialty, Google Professional Cloud Architect, Certified Kubernetes Administrator, SRE Foundation or Practitioner, and observability platform certifications.
  • Must obtain UK Security Check clearance and have continuously resided in the United Kingdom for the last five years, subject to other eligibility criteria.

Benefits

  • Hybrid working with London, Manchester, or Glasgow as an office base and flexible working arrangements.
  • Flexible benefits options, wellbeing support including Mental Health Champions, and access to Thrive and Peppy wellbeing apps.
  • Training and certification opportunities through internal and partner-led programmes from AWS, Google, and Microsoft.
  • Formal training in management consulting and client delivery, plus hands-on exposure to high-profile transformation engagements.
  • Access to Les Fontaines training facilities, monthly showcases, leadership connect sessions, team drinks, and team away days.
  • Opportunities for business development, internal contribution, practice development, whitepapers, offering development, and career learning.
  • The role may require full flexibility regarding assignment location and periods away from home at short notice.

Categories

DevOpsSite Reliability
Capgemini

About Capgemini

10,000+ employees

Capgemini is an AI-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organizations and make it real with AI, technology and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end-to-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations. The Group reported 2025 global revenues of €22.5 billion. Make it real | www.capgemini.com

Contact me