AI value is real, proof is rare, and cost is now very visible. Chris Davies, our client and partner engagement principal, opened our webinar with that line last month. Technology leaders have stopped asking whether AI is worth the money. They now have to prove it, on a ledger where the spend is itemised to the dollar and the return is not itemised at all.
Chris co-hosted AI at scale: cost, context and delivery performance with principal consultant Serge Matheson and software engineering consultant Rahul Samaranayake. We built it around a problem we keep hearing: organisations that have moved past experimentation without building a repeatable way to deliver. Watch the webinar on demand here.
Adoption is rising, token spend with it, engineers are visibly faster, and the release cadence looks much like it did last year. Chris hears some version of this almost daily. He sits across a portfolio of engagements rather than inside any one of them, which is part of why the pattern is visible to him at all.
Afterwards I sat down with him to push on why so few organisations manage to make any of this repeatable. His answer is about how work moves through a team: who hands off to whom, how many times a requirement bounces back before anyone can build it, where the work sits waiting. No model debate or licence purchase touches any of that.
Why the tool question keeps coming back
Chris traces a shift in the question organisations bring him. Two years ago it was “are we doing enough with AI?” Then, “what tools and models are you using?” Now it is “how do we demonstrate value, without cost, inconsistency and governance overhead growing faster than the value?”
His answer has not changed in three years.
“The problem is not technology. It’s people and process.”
The tooling question keeps coming back at him anyway. Which model, which platform, which licence.
Model choice is legible. It has a price, a benchmark and an owner. “Change how your teams work” has none of those, and it asks a leader to spend political capital they may not have on a result they cannot schedule.
Chris pushes it further. Pick the wrong model and you swap it out. The decision is reversible, and some of the blame lands on the vendor. Change how your teams work and get it wrong, and the blame is yours alone. Nobody issues an RFP for “change how code review works.”
Few of the people asking think it is the real question. It is the one their organisation has machinery for answering. Which means the real answer has to be startable, so I asked Chris what the first month actually looks like.
His answer is a two-week sprint against a single bottleneck in a single team. The room is deliberately small: an engineering or product leader who owns the outcome and ideally has the authority to change it, a subject matter expert, and someone who owns delivery.
The first decision is which constraint to attack, and it is almost never code generation. On a current engagement it is the requirements handoff. How many cycles of back and forth it takes to get a feature from a solution document to something a team can actually build.
The second decision is what to measure before anything changes. Cycle time per ticket, from pickup to production. Deployment frequency. How many iterations a requirement goes through before it closes.
“Baseline first, or you will never be able to say what the AI did.”
By week two the goal is modest and specific: one workflow instrumented end to end, a baseline that did not exist before, and a team working differently in a way that is written down.
That shape repeats across the current client base. One organisation is consolidating several platforms into one, because teams were working in silos with no central rigour. Another, in financial services, has regulatory obligations that have to be built in from the start rather than added at review. Different sectors, the same shape underneath.
What’s getting measured, and what should be
On the webinar, Serge described token leaderboards as “a relic of when people were measuring AI enthusiasm, not AI productivity.” Rahul was blunter: treating token spend as a goal in itself is “asinine.”
Chris has seen the same. Teams tracking usage and calling it progress. Spend that used to sit inside an experiment budget now arriving as a line item, with no link back to anything shipped.
The unit that matters, as Chris put it, is the loop end to end. What did it cost to ship a feature. How long did it take. Did it survive contact with production. The signals are the familiar ones: cycle time, pull request size, change failure rate, defect density.
Those four are the signals of good software engineering and have been for more than a decade. None of that stops being true because a model wrote the first draft. The volume of code arriving for review has gone up sharply, and the signals that tell you whether it holds up are the ones we already had.
So I asked Chris what changes, if the measures themselves don’t.
His answer was that the numbers only mean something against a goal someone has committed to. Serge’s example from the webinar was an internationalised version of the product live in a new country in four weeks. That is the benchmark. Are you trending towards it?
The goals on Chris’s current engagements are what you would expect of any technology project. A deadline in October to show something to vendors. A software platform moving from quarterly releases to fortnightly. A retailer that wants 30 per cent of tickets resolved on first contact, ramping to 80.
“None of these mention AI, and all of them are answerable.”
Why individual speed doesn’t spread
The webinar described a spectrum, and the thing that moves along it is where the human sits. At one end, an engineer working line by line with a copilot, reviewing every suggestion as it appears. At the other, agents running overnight against a spec and opening pull requests by morning, with people reviewing at the edges rather than inside the work.
That spectrum runs inside organisations. Most of the ones we work with have a handful of people operating near the far end and a delivery system that has not moved. Those people get much faster. Nothing around them changes, so nothing downstream does either. They were never the constraint.
A product company can move the whole system. Steve Bartlett at Programa, owns the decisions that set the pace: which models are approved, which systems connect to what, how review works. In a large enterprise those same decisions accumulate as policy, and that policy sets the ceiling. Meanwhile the tool spend keeps climbing, because adding a licence needs nobody’s permission and changing how a team works needs everybody’s.
DiUS engineer James Menzies, had eight-hour tasks down to 30 minutes, until he hit a system that did not allow agent access.
“Individual contributors can do great things,” Chris said. “That doesn’t necessarily help the organisation.”
I asked him what breaks the deadlock, given the person who could authorise agent access usually sits three levels away from the person who needs it.
His answer on the webinar was a precedent. Engineers once wanted to deploy to production several times a day, and the organisation said no for the same reason it says no now. There was no basis for confidence. What changed it was not persuasion. It was tests, rollbacks and a pipeline.
“The brakes don’t make a car go slow, but they make going fast a lot safer.”
Serge put the same thing in terms of the harness. Sandboxed permissions, so an agent cannot reach past what it has been handed. A checkable output at the end of each phase, so a person can inspect the work before the next one starts. Telemetry pooled across every run, not only the one someone happens to be watching. Build that and the permission follows, the same way continuous delivery earned teams the right to ship daily.
We have done this before
“It’s a systems thinking problem,” Chris said near the end of our conversation.
That holds across everything we covered. Measurement tied to delivery outcomes. Access and governance decided on purpose. Controls built before capability. None of it is complicated in principle. The difficulty is doing it consistently, across a team, at the pace the business is asking for.
The part I keep coming back to is that this is not the first time. We have already twice changed how software gets built. Agile, then cloud. Both were argued about for years. Both went through an early phase where the tooling question stood in for the operating model question: which CI server, long before anyone moved to trunk-based development. Which region and instance type, long before anyone changed who was allowed to deploy. Both eventually landed, and where they landed properly it was because someone treated it as a change to how people worked.
The pace is different this time. The pattern looks the same to me.