Introducing Deliberate AI: The Method That Gets Enterprise AI into Production

First in a blog series that covers Deliberate AI, our proven, production hardened enterprise AI transformation method.

Marchmont Group is a fictional business, drawn from our AI transformation work across financial services, energy and utilities, FMCG and eCommerce with global enterprise organisations. 


Claire Ashworth was given the AI portfolio on a Tuesday afternoon, in a conversation with her line manager that lasted roughly eleven minutes. Her predecessor had resigned the previous Thursday and was not working his notice, which left a handover consisting of a shared drive folder and a standing offer to answer questions by email. Claire never took him up on it. By the time she had spent an evening with the folder, she understood that the questions it raised were not ones he would have been able to answer.

The folder contained a spreadsheet, a strategy and a build, and between them they told a story that will be familiar to anyone who has inherited an enterprise AI programme at a comparable stage. The spreadsheet ran to one hundred and forty-two rows, one per candidate use case, colour coded green, amber and red by someone who had not left a legend. 

Sixty-one of the rows carried the name of a managing director who had left the group in April in the column headed Sponsor, which meant that nearly half the portfolio no longer had anyone accountable for it. Every row carried a figure in the column headed Estimated Benefit, and those figures had been populated in a single afternoon eighteen months earlier, never revisited, and quoted in board papers ever since as though they were settled fact. 

The strategy was a hundred and eighty slides produced by Vance and Court over eleven weeks, accompanied by an operating model design featuring four proposed centres of excellence and a three-horizon roadmap. Claire read the whole thing over a weekend and found very little she disagreed with, which made it all the more striking that the document properties showed it had last been opened five months before, by her. 

The build was an AI platform engagement with Corvin Systems, working to a specification signed off the previous spring. The programme was eleven weeks behind schedule, the weekly status report was green on four of six measures, and when Claire asked for access to something she could log into and use, it took nine days to establish that no such environment existed outside a demonstration instance that was rebuilt before each steering committee.

On Thursday of that same week, Marchmont's chief executive, Duncan Hale, asked Claire what he could tell the board in March. Eighteen months earlier he had told the market that the group would be AI-native within two years, and sitting underneath that sentence was a cost-to-income commitment that had not moved since the day it was made.

Claire went back to the spreadsheet that evening and understood something that it took her another month to say out loud. She had ideas, because there were one hundred and forty-two of them. She had a strategy, because there were a hundred and eighty slides of it. She had engineering capacity, because Corvin had forty people on the account. She had all three of those things and nothing whatsoever running in production, and no defensible basis on which to decide which of the one hundred and forty-two things to start on Monday, and no amount of additional strategy, additional ideas or additional engineering capacity was going to change that, because the problem was not a shortage of any of those things.

The inputs existed. 

The outcomes did not. 

WeBuild-AI arrived six weeks later.

The two constant AI transformation blockers we see

Almost every large organisation we work with is stuck between the same two blockers, and the reason the situation is so difficult to escape from the inside is that those two blockers present themselves, to the people living with them, as each other's cure.

The first is strategy that never reaches execution. The artefacts are familiar: research papers, roadmaps, target operating models, maturity assessments and prioritisation matrices, produced to a high standard by capable people and peaking in value on the day they are presented. There is nothing wrong with any of this work in itself, and we say that without qualification, because the strategies we encounter are usually sound. The difficulty is that they carry no mechanism for becoming something a user can touch. A strategy document describes a destination and a rationale for travelling there, but it does not describe the road surface, the vehicle or the driving conditions, and when execution does not follow within a reasonable window, the natural institutional conclusion is that the analysis was insufficient. The organisation commissions more of it, the folder grows thicker, and the distance between what has been described and what has been built remains exactly where it was.

The second is technology that never reaches production. The artefacts here are equally familiar: pilots, proofs of concept, hackathon outputs and demonstration environments, some of them genuinely impressive in what they show, none of them carrying the security posture, the data governance, the observability or the operational ownership that would be required to run against real users and real money. The diagnostic is the one Claire applied without meaning to. She asked for a login, and the length of time it took to answer her told her everything the status report had not. At Marchmont it took nine days to establish that the only thing anyone could show her was a demonstration instance rebuilt before each steering committee, and throughout those nine days the programme continued to report green on four of six measures, because the measures it was tracking had been defined around activity rather than outcome. When builds of this kind stall, the natural institutional conclusion is that the strategy was unclear, and so the organisation commissions a strategy, which brings the cycle neatly back to the first failure mode.

Marchmont had been running both loops simultaneously for two years, each one reinforcing the other with considerable energy and sincerity, producing a volume of activity that would satisfy any casual observer and no production capability whatsoever.

The gap is a deliberate route to production, not a plan

What sits between those two failure modes is not a better plan and not a better proof of concept. It is a dependable, repeatable route from an idea to a running AI system, one that can be walked again and again within a regulated estate, using the controls the organisation already has, without requiring each new build to reinvent the path from scratch.

That is a narrower thing than a transformation programme and a considerably more demanding one, because a route has to deliver a result every time it is followed, not simply convince a steering committee on the day it is presented. It has to work inside the constraints that actually govern the organisation: the legacy platforms that are not going anywhere, the reference data that everyone knows is imperfect, the model risk function that requires evidence before it will approve anything, and the reasonable scepticism of the people whose working day is about to change. The test of a route is whether it puts a working AI system into production and whether it can do so repeatedly, because that is a materially harder thing to achieve than a convincing strategy or an impressive demonstration, and it is the only thing that counts.

This is why we have created Deliberate AI.

Introducing Deliberate AI

Deliberate AI is our attempt to write that route down plainly, drawing on the lessons of delivering AI into large, global, regulated enterprises. It has three parts, and each one addresses a different dimension of the problem that Claire was living with when we first met her.

The first is our Method, which is the route from discovery to production, phased across five stages. 

Map, establishes what the organisation actually runs on, which is frequently not what the organisation chart or the process documentation suggests. 

Shape, resolves the accumulated demand into the capabilities that sit underneath it, distinguishing between genuinely distinct requirements and surface variations of the same need. 

Ship, builds the first of those capabilities in production, against real data and real users, as a Lighthouse Project that demonstrates value before anything is accelerated.

Extend builds the next wave of capabilities from the cross-cutting foundations that the Lighthouse leaves behind, so that each subsequent build is faster and cheaper than the one before it. 

Scale takes those capabilities across the organisation. The purpose of the sequence is not to slow things down but to ensure that a hundred ideas become the right builds in the right order, rather than a queue serviced by whoever asked loudest or most recently.

The second is the Toolkit, which is the eight engineering layers on which every solution is assembled and operated. These run from infrastructure and operations at the base, through security, evaluation, data, models, agents and orchestration, up to the experience layer that a customer or employee actually touches. 

Each layer is backed by opinionated accelerators that converge open source, cloud-native and frontier AI capabilities into a production-ready foundation, meaning that the same toolkit runs on AWS, Azure, Google Cloud or private cloud infrastructure, the last of these supported through our strategic relationship with NVIDIA. 

The result is an engineering spine that prevents every new build from starting with a blank repository, and that gives the client genuine portability rather than a dependency on a single provider.

The third is the Motion, which is how we work. Our engineers deploy into the client's teams and the client's cloud, building alongside their people so that the capability, the knowledge and the operational ownership remain when we step back. 

We have worked this way since the inception of the business, and it is worth noting that the wider industry conversation around forward deployed engineering has since converged on the same model. The difference is that we are not adopting it now because it has a name. We built the firm around it, and the Deliberate AI method is the codification of what that way of working has taught us over several years of applied AI delivery into regulated enterprises.

The five principles Deliberate AI is built on

We have refined these principles since our inception, and every one of them exists because we learned, usually the hard way, what happens when it is absent. They are not theoretical. They are the commitments we make to ourselves and to our clients before an engagement begins, and they govern how we work from the first week to the last.

Targeted. We designed Deliberate AI to reach production AI solutions quickly rather than to be complete on paper, which means it is not all-encompassing, documentation-heavy or rigid. The first phase produces a queryable ontology of the operating estate in a matter of days, not weeks, and the test that model has to pass is whether it can tell a leader in Claire's position which build to start on Monday and why. That is a deliberately narrow test. We set the bar there because we have seen too many discovery phases produce thorough, impressive analysis that could not answer the most basic operational question a sponsor will ask, and a phase that cannot answer it has not earned the time it consumed.

Prescriptive. Our method offers a template and a sequence rather than a menu, and we are unapologetic about that, because the organisations we work with do not need more optionality. Claire already had a hundred and eighty slides setting out her options in considerable detail. What she did not have was a recommendation with an order attached to it, delivered with enough conviction, clarity and enough evidence to act on immediately. We have watched other engagements respond to that situation by offering a further set of choices, and the result is always the same: it adds to the problem it was brought in to solve.

Integrated. We account for people, process and technology together in every engagement, because AI programmes that treat any one of those dimensions in isolation reproduce exactly the failure modes described above. A classifier that meets every accuracy threshold and that the operations team decline to use has failed, and it fails quietly, surfacing months after the technical work was signed off in an adoption report that nobody quite knows what to do with. We learned early on that treating adoption as a downstream communications exercise rather than an upstream design constraint is one of the most reliable ways to waste an AI investment, and we structured the method to prevent it.

Proven. We insist that a lighthouse project is used to prove value in production, in the client's own environment, before anything is accelerated. Proof precedes scale in every case, without exception, because the alternative is to invest at scale on the basis of a demonstration that cannot tell you whether the system will work under real conditions. In practice this means that the first capability goes live against real data and real users while it is still visibly imperfect, and we have found that the feedback this generates is precisely what makes the second and third capabilities materially better than the first. The demonstration instance that Claire spent nine days looking for, and that told nobody anything about whether the system worked, is what happens when an organisation attempts to prove value without putting anything into production.

Instrumented. We obsess about business value and outcomes, and we measure them from day one, because if we do not then six to nine months later someone senior will ask where the value is and nobody will have a credible answer. We require the baseline to be agreed before a single line of code is written, and we track quality, cost and value continuously as the build progresses. This can feel unnerving for our clients, because establishing a measurable baseline means putting a number on something that will be held to account, and that carries real personal exposure if it does not come off. We understand that. What we give them in return is the data and the evidence that ensures the wider business understands why the change is pivotal, which means we can unblock at every turn rather than waiting for a quarterly review to discover that momentum has stalled.

Questions worth asking in your own organisation

If any of the above was uncomfortably familiar, four questions will tell you where you actually stand, and the answers are worth having before the next board paper is written to update on your AI strategy.

The first is how many of your candidate use cases are genuinely in production, being used by people who would notice and complain if you switched them off, rather than merely deployed, demonstrated or sitting in a pilot that has not yet earned a decision. In our experience, the gap between the number of use cases on the roadmap and the number that meet this test is the single most revealing metric in any enterprise AI programme.

The second is when your current AI strategy last changed as a result of something you learned by building. A strategy that is never updated by delivery feedback is a document rather than a plan, and the distance between those two things grows wider with every month the feedback loop remains open. If the strategy reads exactly as it did on the day it was presented, the organisation is navigating by a map it drew before it set out, and that is a map whose accuracy is declining by the week.

The third is how long it would take, if you asked today, to be shown something running, and whether what you were shown would be an AI live system serving real users or a demonstration prepared for the occasion. The answer to this question tells you more about the true state of a programme than any status report or steering committee pack, because the difference between a live system and a demonstration is the difference between evidence and argument.

The fourth is what the baseline for quality, cost and value was before your last AI build started, and who agreed to it. If the answer is that no baseline was set, then the organisation has no way of knowing whether the build improved anything, no way of comparing alternatives, and no way of defending its investment to the people who approved it. A benefit case that is constructed after the work is complete, using whichever measures happen to look favourable, is not a benefit case. It is a post-hoc justification, and most stakeholders can tell the difference.

What comes next in this series

The next post in this series covers what happened when WeBuild-AI arrived at Marchmont, and why the first thing we asked for was repository access and a desk in settlement operations rather than a discovery workshop.

It sets out how we use forward deployed engineering as a motion for change, embedding alongside the client's own teams from day one, and how we apply AI to map the processes, gaps and challenges across the operating estate so we can understand the current state and optimise it with and for AI, in future phases of delivery.

Next
Next

From disparate spreadsheets to a secure, scalable AI platform in 6 months