specshop.dev / Journal / Industry take
Industry take 7 min read

You Bought the Contractor. You Never Wrote the Brief.

Buying Copilot licences transfers capability without transferring instruction. Two government trials, MIT, Gartner and BCG all circle the same missing artefact — the brief nobody ever wrote down.

Janaka Ediriweera
CEO & Co-founder · specshop.dev

Buying Copilot licences for your organisation is the equivalent of hiring an exceptionally fast contractor, issuing them a security badge, and never writing the brief.

The contractor will be busy. They will produce things. Some of those things will be useful. None of it will change how your organisation works, because nobody ever told them what the work is.

This is not an argument against Copilot. Copilot does what it says. The argument is narrower and more uncomfortable: a licence transfers capability without transferring instruction, and most organisations have never written down what they actually do.

§

Two governments, one tool, opposite conclusions

In late 2024 the UK Department for Business and Trade handed out a thousand Copilot licences and evaluated the result over three months. Around three hundred participants consented to having their data analysed. The findings were mixed in a very specific way. Writing emails, summarising meetings and producing report summaries got faster. Data analysis in Excel got slower and less accurate. Slide production was quicker on average but the output needed correcting. On the headline question — whether any of this added up to a measurable productivity improvement — the department found no conclusive evidence.

Australia’s Digital Transformation Agency ran its own trial and reported something that sounds like the opposite. Roughly seven in ten participants agreed Copilot improved the speed at which they completed tasks. Six in ten said it lifted the quality of their work. Around two thirds of managers saw a positive effect on their teams.

The trial What it measured The headline
UK DBT measured output ~1,000 licences, three months, ~300 consenting to data analysis; observed change in actual output Some tasks faster, Excel slower — no conclusive productivity gain
Australia DTA measured feeling Whole-of-government trial; participants self-assessing the impact on their own work ~70% felt faster, ~60% felt higher quality, ~⅔ of managers positive

The two results are not actually in conflict, and the DTA said so itself. Its evaluation relied on participants self-assessing the impact, and the agency explicitly flagged the risk that its productivity estimates might not reflect Copilot’s actual productivity impact.

One government asked people how it felt. The other tried to measure what changed. That gap — between the felt experience of assistance and the observable change in output — is where several billion dollars of enterprise AI spend is currently sitting.

§

The specification never existed

The instinct is to read those results as a verdict on the model. It isn’t. It’s a verdict on the brief.

An AI assistant plugged into a knowledge worker’s day inherits that day exactly as it is. It inherits the meeting nobody can explain the purpose of. It inherits the report that three people read and none of them act on. It inherits the approval step that exists because someone was burned in 2019. Point a very capable tool at an unspecified process and you get the same unspecified process, executed with more fluency.

Most organisations cannot describe how their own work gets done. Not because they are badly run, but because the description was never required. Process lives in people’s heads, in habit, in the tacit knowledge of whoever has been there longest. It functions perfectly well as long as the only things executing it are humans, who are extraordinarily good at filling gaps nobody wrote down.

The moment you introduce something that executes literally, that tolerance disappears. Vagueness stops being a cultural quirk and becomes a failure mode. The specification you never wrote turns out to have been the product all along.

§

Your staff already solved this. For themselves.

Here is the part that should worry a CFO more than the idle licences.

MIT’s Project NANDA study of enterprise generative AI put the proportion of pilots delivering no measurable impact on profit and loss at ninety-five percent. That number is contested and worth treating as directional rather than precise. But the study’s second finding is harder to argue with and much less discussed: while sanctioned pilots stall, employees across the overwhelming majority of firms are using personal AI tools anyway, off the books, to get their own work done faster.

95%
MIT Project NANDA put the share of enterprise GenAI pilots with no measurable P&L impact at ninety-five percent. Treat the figure as directional. The harder-to-argue finding sits next to it: the value is being extracted — privately, off the books, by employees using personal tools — while the sanctioned pilot shows nothing.

Read that carefully. The value is being extracted. It is simply being extracted privately.

An individual can absorb a new tool into their workflow in an afternoon, because an individual’s workflow is a specification of one, held entirely in their own head, and they are free to rewrite it. An organisation cannot, because its workflow is a negotiated settlement between dozens of people, systems and assumptions that nobody has authority to unilaterally redraft.

So the shortcut stays personal. The employee gets home earlier. The P&L never finds out. And the licence you paid for formalised something people were doing anyway rather than changing what the organisation is capable of.

§

Gartner’s maturity model is a specification maturity model wearing a suit

Gartner assesses AI maturity across strategy, data, technology, governance, talent and business value, sorting organisations into five stages from foundational experimentation through to transformational. It is worth noticing what is absent from that list. Licences purchased is not a dimension. Seats deployed is not a dimension. Tool count does not appear anywhere.

Every dimension that does appear is a specification problem in disguise. Governance is a written specification of what is permitted. Data quality is a specification of what a record means. An operating model is a specification of who decides what. Business value measurement is a specification of what success looks like before you go looking for it.

Gartner’s own survey data, drawn from over four hundred organisations across six countries, shows what happens when those specifications exist. High-maturity organisations keep AI initiatives running in production for three years or more at more than twice the rate of low-maturity ones. Business units trust and are ready to adopt new AI solutions in well over half of high-maturity organisations, against roughly one in seven of the low-maturity group. Around two thirds of the mature group run actual financial and ROI analysis on their AI work.

None of that is a technology gap. It is entirely a documentation and decision-rights gap. The mature organisation didn’t buy a better model.

§

The cheapest line item

Boston Consulting Group’s guidance to its clients is that ten percent of the effort in an AI transformation should go to algorithms, twenty percent to technology and data, and the remaining seventy percent to people and processes.

70%
BCG’s 10–20–70 rule: ten percent of an AI transformation is the algorithm, twenty percent the technology and data, and seventy percent people and processes. A licence purchase lives inside the first ten percent — the smallest, most easily completed part, which is exactly why it gets finished first and mistaken for progress.

A licence purchase is a transaction inside the first ten percent. It is, by BCG’s own arithmetic, the smallest and most easily completed part of the job — which is precisely why it is the part everyone completes first and then mistakes for progress. The procurement was fast, the invoice was real, and the organisational change was zero.

The other seventy percent is the brief. It is the unglamorous work of writing down how the work actually gets done, deciding which parts of it should stop existing, and specifying clearly enough that something which executes literally can be trusted to execute it.

Nobody wants to buy that. It doesn’t have a per-seat price and it can’t be signed off in a quarter.

The licence was the cheapest thing you bought. The expensive part is admitting no one ever wrote down how the work is done.