Tag: Artificial Intelligence

  • AI Is Changing White-Collar Work. Measuring It Is the Hard Part

    Organisations deploying AI tools across knowledge work are discovering an awkward problem: they cannot reliably tell whether it is working. This is not primarily a failure of the technology. It is a failure of measurement infrastructure that predates the technology by decades, and which AI adoption has merely exposed.

    Productivity has a definition, and it is not “speed”

    Productivity is output per unit of input. For manufacturing this is tractable: count units produced, count hours worked, divide. The output is physical, countable and homogeneous.

    Knowledge work breaks every one of those conditions. What is the output of a lawyer, an analyst, a designer, a manager? Documents produced is a measure of activity, not value — a lawyer producing twice the contracts of similar quality has doubled output; one producing twice the pages has not.

    Because genuine output is hard to measure, organisations substitute proxies: hours logged, tickets closed, lines of code, documents drafted, meetings held. These proxies were weak measures before AI. They are actively misleading now, because AI tools improve exactly the proxies while leaving the underlying question untouched. A team can double its document output and produce no additional value whatsoever.

    The task-to-firm gap

    The most consistent finding in research on AI and work is that measured gains shrink as the unit of analysis widens.

    At the level of a discrete, well-specified task — draft this summary, write this function, translate this document — controlled studies have generally found substantial time savings, frequently with the largest relative gains among less experienced workers, for whom the tool substitutes partially for expertise.

    At the level of the firm, these gains have been considerably harder to detect in financial results. The gap has several sources, and none of them are mysterious.

    Saved time must be redeployed to be worth anything. If a task that took two hours now takes one, the organisation captures value only if that hour is used for something productive. Often it is absorbed into slack, longer meetings, or additional revisions of work that was already adequate.

    Bottlenecks move. Accelerating drafting does not accelerate a process gated by legal review, client response or a monthly approvals meeting. The constraint relocates rather than disappearing, and total throughput barely shifts.

    Verification costs are real and frequently uncounted. Output that must be checked for accuracy carries a review burden. Where checking is nearly as expensive as producing — as it often is for factual, legal or numerical content — net savings can approach zero even when drafting time falls sharply. Studies measuring generation time without measuring verification time systematically overstate gains.

    Quality changes are not captured. If output quality improves, productivity gains are understated. If quality degrades in ways that surface later — errors caught downstream, rework, reputational cost — gains are overstated. Most measurement systems capture neither.

    An old pattern

    This is a recognisable historical shape. Robert Solow’s 1987 observation that the computer age was visible everywhere except the productivity statistics described the same phenomenon for information technology, and it took years before measured productivity growth clearly reflected computing investment.

    The explanation developed since — most associated with Erik Brynjolfsson and co-authors, and often called the productivity J-curve — is that general-purpose technologies require large complementary investments in intangibles: reorganised processes, retrained staff, restructured workflows, new management practice. Those investments are costly and are typically expensed rather than capitalised. During the transition, measured productivity can appear worse, because the costs are recorded immediately while the benefits accrue later and are partly invisible to national accounts.

    If that pattern holds, the current difficulty in measuring AI’s effect is expected rather than evidence of failure — and equally, it is not evidence of success. It is what an ambiguous transition legitimately looks like.

    What organisations are actually measuring

    In practice, most AI measurement programmes track adoption rather than outcomes: licences issued, weekly active users, queries submitted, self-reported time saved.

    These are usage metrics. They establish that a tool is being used, not that it is creating value. Self-reported time savings are particularly unreliable — respondents estimate against a counterfactual they never observed, and are subject to well-documented optimism when reporting on tools they have chosen to adopt.

    Approaches that produce usable evidence

    Several methods yield defensible answers, and all of them require more discipline than a dashboard.

    • Staggered rollout with a control group. Grant access to part of the organisation first and compare outcomes against a comparable group without access. This is the closest most firms can get to a controlled experiment, and it is administratively straightforward if planned before deployment rather than after.
    • Measure end-to-end cycle time, not task time. Track the interval from work initiation to completed, accepted delivery. This captures bottleneck relocation and verification burden, both of which task-level timing misses.
    • Instrument quality explicitly. Error rates, rework frequency, downstream complaints and revision counts. Without a quality measure, any throughput gain is uninterpretable.
    • Track where saved time goes. If capacity is freed, establish what it was redeployed to. Unredeployed capacity is not a productivity gain.
    • Separate experience levels. Effects have consistently differed between novice and expert workers. Blended averages conceal both.

    The measurement trap to avoid

    The strongest temptation is to adopt whichever metric moves most, since it produces the most persuasive internal narrative. This is Goodhart’s law waiting to operate: once a proxy becomes a target, it stops measuring what it was chosen to represent.

    An organisation that rewards teams for AI-attributed output volume will reliably get more output volume. Whether it gets more value is a separate question that the metric has been structurally designed not to answer.

    The reasonable position

    Both confident narratives — that AI is transforming white-collar productivity, and that it is delivering nothing — currently outrun the available evidence. Task-level gains are well documented. Firm-level gains are harder to detect, for reasons that are understood and that have precedent.

    The organisations that will know the answer first are those that built measurement into deployment rather than attempting to reconstruct it afterwards from usage logs.

    Related reading

    For a comparable case of cost assumptions outrunning evidence, see why some companies are moving workloads out of the cloud.

    This article is general information and journalism. See our Editorial Policy.