hirly

Apply with hirly

AI Platform & Site Reliability Engineering Managing Consultant

Capgemini Invent · Glasgow, Manchester, London, GB

Upload your resume to see how well you match this job — free, in seconds, no account needed.

Your resume is used only to score it against this job. If you don't create an account, it is deleted within 24 hours.

About Capgemini

At Capgemini Invent, we believe difference drives change. As inventive transformation consultants, we blend our strategic, creative and scientific capabilities, collaborating closely with clients to deliver cutting-edge solutions. Join us to drive transformation tailored to our client's challenges of today and tomorrow. Informed and validated by science and data. Superpowered by creativity and design. All underpinned by technology created with purpose.

Your Role

As an AI Platform & Site Reliability Engineering Managing Consultant, you will help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services. You will work with technology, engineering, operations and business leaders to establish the platforms, operating models, governance and reliability practices required to run AI-enabled services safely, effectively and at scale. Acting as a trusted advisor to senior stakeholders, you will shape client strategy while leading delivery teams and helping grow our AI Platform & Reliability Engineering capability. This will include: • AI Platform Strategy & Architecture: Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite and operational requirements. • AI Platform Engineering & LLMOps: Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. • Reliability Engineering & SRE: Establish SRE practices including SLIs, SLOs, error budgets, capacity planning, resilience engineering and reliability governance. Help clients shift from reactive operations to data-driven reliability management while balancing reliability, innovation and delivery velocity. • Observability & Operational Intelligence: Define observability strategies across applications, platforms and AI workloads using metrics, logs, traces and telemetry. Establish operational insight models that support proactive decision-making and enable advanced capabilities including anomaly detection, event intelligence, noise reduction and predictive operational analytics. • AI Operations & Service Reliability: Apply reliability engineering principles to AI-enabled services, monitoring AI-specific failure modes such as data quality degradation, hallucination patterns, token consumption, agent reliability and model performance drift. Implement controls, feedback loops and automated guardrails to ensure AI services remain secure, trusted and cost-effective. • Responsible AI & Platform Governance: Design and embed AI governance frameworks, model risk controls, compliance measures and responsible AI practices that address security, regulatory and ethical requirements while supporting innovation and adoption at scale. • Automation & Operational Efficiency: Identify opportunities to reduce operational complexity and toil through engineering-led automation, intelligent workflows and AI-enhanced operational practices. Help clients improve scalability, consistency and operational performance while reducing manual effort. • Client Advisory & Transformation Leadership: Act as a trusted advisor to CIO, CTO, CDO and Engineering leadership stakeholders, shaping platform strategies, operating models, vendor selections and transformation roadmaps. Lead consulting teams and workstreams from assessment and strategy through implementation and scale-up.

Your Profile

Essential Experience • Proven experience designing, delivering and operating cloud-native, platform engineering, AI platform or reliability engineering solutions within complex enterprise environments. • Strong understanding of AI platform architectures including LLMOps, MLOps, agentic AI frameworks, model lifecycle management and AI operational controls. • Experience establishing and scaling SRE practices including observability, SLIs, SLOs, error budgets, incident management and reliability engineering. • Strong understanding of AI governance, responsible AI, regulatory requirements and model risk management. • Experience implementing observability strategies using modern monitoring, telemetry and operational analytics platforms. • Demonstrated ability to advise senior stakeholders and lead multidisciplinary transformation programmes. • Experience working across hyperscaler ecosystems including Azure, AWS and Google Cloud Platform. • Proven ability to balance business outcomes, user needs, engineering constraints and operational requirements when shaping platform strategies. Desirable Experience • Experience with AI observability, model monitoring, AI governance tooling or AI platform operations. • Experience of platform engineering, DevSecOps, automation and Infrastructure-as-Code practices. • Experience developing propositions, leading bids and supporting business growth activities. • Active participation in AI, SRE, platform engineering or cloud communities. Certifications (Desirable) • Azure AI Engineer Associate • Azure Solutions Architect Expert • AWS Machine Learning Specialty • Google Professional Cloud Architect • Certified Kubernetes Administrator (CKA) • SRE Foundation or SRE Practitioner • Relevant observability platform certifications (Datadog, Dynatrace, Splunk etc.) Security Check (SC) Clearance To be successfully appointed to this role, it is a requirement to obtain Security Check (SC) clearance. ( To obtain SC clearance, the successful applicant must have resided continuously within the United Kingdom for the last 5 years, along with other criteria and requirements. Throughout the recruitment process, you will be asked questions about your security clearance eligibility such as, but not limited to, country of residence and nationality. Some posts are restricted to sole UK Nationals for security reasons; therefore, you may be asked about your citizenship in the application process.

What You'll Love About Working Here

• Client engagements give you the opportunity to work with our leadership and experienced consulting management, where you can learn from them, challenge them, and accelerate your hands-on experience, delivery capability and industry insights. • You’ll learn how we write compelling client propositions, structure, and lead high-profile transformation, and gain hands-on exposure to leading technologies, often taking an idea from a concept to a vision, to strategy and then execution. • Our consultants are formally trained from industry experts on management consulting and client delivery. • We provide a host of opportunities for learning and certification through internal and partner led programmes from AWS, Google and Microsoft. • Les Fontaines: Capgemini Invent has a unique training environment just outside of Paris, where we can immerse ourselves in thought-leadership, share knowledge and build capabilities which will help us and our clients to succeed. • We hold monthly showcases of our initiatives, sharing knowledge and showing off how the power of technology is impacting our clients. • There are many opportunities like monthly team drinks to con