Category: Analysis

Longer-form interpretation of business and economic developments.

  • AI Is Changing White-Collar Work. Measuring It Is the Hard Part

    Organisations deploying AI tools across knowledge work are discovering an awkward problem: they cannot reliably tell whether it is working. This is not primarily a failure of the technology. It is a failure of measurement infrastructure that predates the technology by decades, and which AI adoption has merely exposed.

    Productivity has a definition, and it is not “speed”

    Productivity is output per unit of input. For manufacturing this is tractable: count units produced, count hours worked, divide. The output is physical, countable and homogeneous.

    Knowledge work breaks every one of those conditions. What is the output of a lawyer, an analyst, a designer, a manager? Documents produced is a measure of activity, not value — a lawyer producing twice the contracts of similar quality has doubled output; one producing twice the pages has not.

    Because genuine output is hard to measure, organisations substitute proxies: hours logged, tickets closed, lines of code, documents drafted, meetings held. These proxies were weak measures before AI. They are actively misleading now, because AI tools improve exactly the proxies while leaving the underlying question untouched. A team can double its document output and produce no additional value whatsoever.

    The task-to-firm gap

    The most consistent finding in research on AI and work is that measured gains shrink as the unit of analysis widens.

    At the level of a discrete, well-specified task — draft this summary, write this function, translate this document — controlled studies have generally found substantial time savings, frequently with the largest relative gains among less experienced workers, for whom the tool substitutes partially for expertise.

    At the level of the firm, these gains have been considerably harder to detect in financial results. The gap has several sources, and none of them are mysterious.

    Saved time must be redeployed to be worth anything. If a task that took two hours now takes one, the organisation captures value only if that hour is used for something productive. Often it is absorbed into slack, longer meetings, or additional revisions of work that was already adequate.

    Bottlenecks move. Accelerating drafting does not accelerate a process gated by legal review, client response or a monthly approvals meeting. The constraint relocates rather than disappearing, and total throughput barely shifts.

    Verification costs are real and frequently uncounted. Output that must be checked for accuracy carries a review burden. Where checking is nearly as expensive as producing — as it often is for factual, legal or numerical content — net savings can approach zero even when drafting time falls sharply. Studies measuring generation time without measuring verification time systematically overstate gains.

    Quality changes are not captured. If output quality improves, productivity gains are understated. If quality degrades in ways that surface later — errors caught downstream, rework, reputational cost — gains are overstated. Most measurement systems capture neither.

    An old pattern

    This is a recognisable historical shape. Robert Solow’s 1987 observation that the computer age was visible everywhere except the productivity statistics described the same phenomenon for information technology, and it took years before measured productivity growth clearly reflected computing investment.

    The explanation developed since — most associated with Erik Brynjolfsson and co-authors, and often called the productivity J-curve — is that general-purpose technologies require large complementary investments in intangibles: reorganised processes, retrained staff, restructured workflows, new management practice. Those investments are costly and are typically expensed rather than capitalised. During the transition, measured productivity can appear worse, because the costs are recorded immediately while the benefits accrue later and are partly invisible to national accounts.

    If that pattern holds, the current difficulty in measuring AI’s effect is expected rather than evidence of failure — and equally, it is not evidence of success. It is what an ambiguous transition legitimately looks like.

    What organisations are actually measuring

    In practice, most AI measurement programmes track adoption rather than outcomes: licences issued, weekly active users, queries submitted, self-reported time saved.

    These are usage metrics. They establish that a tool is being used, not that it is creating value. Self-reported time savings are particularly unreliable — respondents estimate against a counterfactual they never observed, and are subject to well-documented optimism when reporting on tools they have chosen to adopt.

    Approaches that produce usable evidence

    Several methods yield defensible answers, and all of them require more discipline than a dashboard.

    • Staggered rollout with a control group. Grant access to part of the organisation first and compare outcomes against a comparable group without access. This is the closest most firms can get to a controlled experiment, and it is administratively straightforward if planned before deployment rather than after.
    • Measure end-to-end cycle time, not task time. Track the interval from work initiation to completed, accepted delivery. This captures bottleneck relocation and verification burden, both of which task-level timing misses.
    • Instrument quality explicitly. Error rates, rework frequency, downstream complaints and revision counts. Without a quality measure, any throughput gain is uninterpretable.
    • Track where saved time goes. If capacity is freed, establish what it was redeployed to. Unredeployed capacity is not a productivity gain.
    • Separate experience levels. Effects have consistently differed between novice and expert workers. Blended averages conceal both.

    The measurement trap to avoid

    The strongest temptation is to adopt whichever metric moves most, since it produces the most persuasive internal narrative. This is Goodhart’s law waiting to operate: once a proxy becomes a target, it stops measuring what it was chosen to represent.

    An organisation that rewards teams for AI-attributed output volume will reliably get more output volume. Whether it gets more value is a separate question that the metric has been structurally designed not to answer.

    The reasonable position

    Both confident narratives — that AI is transforming white-collar productivity, and that it is delivering nothing — currently outrun the available evidence. Task-level gains are well documented. Firm-level gains are harder to detect, for reasons that are understood and that have precedent.

    The organisations that will know the answer first are those that built measurement into deployment rather than attempting to reconstruct it afterwards from usage logs.

    Related reading

    For a comparable case of cost assumptions outrunning evidence, see why some companies are moving workloads out of the cloud.

    This article is general information and journalism. See our Editorial Policy.

  • Index Funds vs Active Management: What the Evidence Actually Supports

    The debate between index tracking and active fund management is often framed as a matter of opinion. A significant part of it is not. One component is arithmetic, true by construction and not dependent on any empirical claim. The remainder is genuinely contested.

    Separating the two makes the argument much easier to follow.

    The arithmetic that is not in dispute

    In 1991 William Sharpe set out an argument now known as the arithmetic of active management. It runs as follows.

    Every share must be held by someone. Divide all holders into passive investors, who hold the market in proportion, and active investors, who do not. Passive holdings, in aggregate, mirror the market — so passive investors collectively earn the market return before costs.

    Since the two groups together own the entire market, and passive investors collectively earn the market return, active investors collectively must also earn the market return before costs. There is no arrangement of ownership in which both groups beat the market, because they jointly are the market.

    Active management is more expensive — research staff, higher management fees, greater trading. Therefore, after costs, the average actively managed dollar must underperform the average passively managed dollar. Necessarily. Not usually, not historically, but as a matter of definition.

    This does not say that no active manager can outperform. It says active management is zero-sum before costs and negative-sum after them: outperformance by one manager is exactly offset by underperformance elsewhere.

    What the empirical record adds

    The arithmetic constrains the average. It says nothing about the distribution — how many managers beat the benchmark, by how much, and whether the same ones do it repeatedly. Those are empirical questions, and they have been studied extensively.

    Long-running scorecards that compare active funds against their benchmarks — most prominently the SPIVA series maintained by S&P Dow Jones Indices, alongside academic work on fund performance persistence — have consistently found the same broad pattern across many markets and asset classes:

    • Over short horizons, a substantial minority of active funds beat their benchmark.
    • As the horizon lengthens to ten or fifteen years, the proportion that outperform falls markedly, commonly to a small minority.
    • Persistence is weak. Funds in the top quartile in one period are not reliably in the top quartile in the next, at rates meaningfully better than chance would produce.

    Two methodological points make these findings stronger than they first appear.

    Survivorship bias. Funds that perform badly are closed or merged away. Studies measuring only funds that still exist systematically overstate active performance. Well-constructed scorecards correct for this, and the correction is substantial — a meaningful share of funds do not survive a fifteen-year window at all.

    Benchmark selection. A fund must be compared to a benchmark matching its actual exposure. A small-cap fund measured against a large-cap index tells you about size exposure, not manager skill.

    Why costs dominate

    Fee differences look trivial annually and are not, because they compound against a growing balance.

    A fee of one percentage point does not reduce a long-run outcome by one percent. It reduces it by roughly one percent of the balance every year, compounded — which over multi-decade horizons removes a large fraction of total accumulated return. The precise figure depends on the return assumption, but the structural point holds under any of them: the drag grows with time and with the size of the balance.

    Costs also extend beyond the headline expense ratio. Trading costs, bid-ask spreads and market impact are borne by the fund and not included in the stated fee. High-turnover strategies incur more of these. In taxable accounts, realised capital gains distributions create a further drag that never appears in any published performance figure.

    The genuine case for active management

    Several arguments survive scrutiny and deserve fair statement.

    Market efficiency varies. Sharpe’s arithmetic holds everywhere, but the dispersion of outcomes does not. In markets with less analyst coverage, poorer disclosure and more constrained participants — smaller companies, some emerging markets, certain credit segments — skill plausibly has more room to operate. Evidence here is more mixed than in large-cap developed equity, where the case against active management is strongest.

    Indices are not neutral. Capitalisation weighting mechanically allocates more capital to companies that have already risen. At concentration extremes an index fund can hold far more in a handful of names than an investor intends. Tracking an index is a deliberate choice about exposure, not an absence of choices.

    Someone must set prices. Index funds free-ride on price discovery performed by active participants. If indexing became universal, prices would stop reflecting information. This is a real theoretical concern, though current indexing levels remain well short of any plausible threshold.

    Objectives differ. Some mandates prioritise downside protection, income stability or specific constraints over benchmark-relative return. Judging such a fund purely on benchmark comparison misses what it was hired to do.

    The identification problem

    The decisive practical difficulty is not whether skilled managers exist. It is whether they can be identified in advance.

    With thousands of funds operating, some will produce excellent long records through chance alone. Distinguishing skill from luck statistically requires far longer track records than most funds possess — and by the time a record is long enough to be convincing, the manager may have retired, the fund may have grown too large to repeat the strategy, or the conditions that suited it may have passed.

    Past performance is the most commonly used selection criterion and among the weakest predictors, which is precisely why regulators require the warning that accompanies it.

    What the evidence supports

    The defensible summary is narrower than either camp’s rhetoric. Costs are the most reliable predictor of relative fund performance available, and they are knowable in advance — unlike returns. The average active dollar underperforms after costs by construction. Outperformance exists but is difficult to identify prospectively and difficult to sustain.

    What follows from that for any individual depends on circumstances, tax position, time horizon and objectives that no article can assess.

    This article is general information and journalism. It is not investment advice, not a recommendation to buy or sell any fund or security, and it does not account for your circumstances. Past performance does not predict future results. Consult a qualified adviser before making investment decisions. See our Editorial Policy.

  • What an Inverted Yield Curve Actually Signals — and What It Doesn’t

    The yield curve inverts when short-term government bonds yield more than long-term ones. In most developed markets this has preceded most recessions of the past half-century, which has earned it a reputation as the most reliable recession indicator available. That reputation is largely deserved and routinely over-applied.

    What the curve is

    Plot the yield on government debt against time to maturity — three months, two years, ten years, thirty — and the resulting line is the yield curve. Its normal shape slopes upward: lending money for longer carries more risk, so it commands more compensation.

    Inversion means that ordering has reversed. Investors accept less annual yield to lock money up for a decade than for two years. On its face this is irrational. Understanding why it is not is the whole point.

    The two components of a long-term yield

    A long-dated yield decomposes into two parts.

    The expectations component is the market’s average expectation of short-term rates over the life of the bond. If you can earn 5% rolling short-term bills for ten years, you will not accept 3% on a ten-year bond — unless you expect short rates to fall well below 5% during that decade.

    The term premium is additional compensation for bearing duration risk: the possibility that rates move against you while your capital is committed. It is normally positive, and it is not directly observable — it must be estimated, which is why credible analysts disagree about its level.

    Inversion occurs when the expectations component falls far enough to overwhelm the term premium. The market is saying, collectively, that it expects short-term rates to be materially lower in the future than they are now.

    Why that implies recession

    Central banks cut rates for essentially one reason: the economy is weakening enough that inflation is no longer the binding concern. An expectation of substantially lower future rates is therefore an expectation of economic deterioration.

    This is the crucial interpretive point, and it is almost always stated backwards in commentary. The inversion is not a cause. It is a summary of what a large, well-capitalised market already believes. The yield curve does not predict recessions in the way a leading indicator does; it aggregates the forecasts of participants with money at stake and displays the result as a single observable number.

    The self-reinforcing mechanism

    There is, however, a genuine causal channel, and it runs through bank profitability.

    Banks fund themselves short — deposits and short-term borrowing — and lend long, in mortgages and commercial loans. Their margin depends on the gap between long and short rates. Invert the curve and that margin compresses.

    Lending becomes less attractive at the margin, so banks tighten standards and ration credit. Credit-dependent borrowers — small businesses especially — find financing harder to obtain. Activity slows. The inversion therefore contributes modestly to the outcome it anticipates.

    Which spread, and why it matters

    “The yield curve” is not one number. Different spreads invert at different times and carry different information.

    • Ten-year minus two-year is the most widely quoted, and the most frequently referenced in market commentary.
    • Ten-year minus three-month has been favoured in a good deal of academic and central bank research on recession forecasting, on the argument that the very short end more directly reflects current policy.
    • Near-term forward spreads, comparing expected short rates a few quarters out against current ones, are preferred by some researchers as a cleaner read on expected policy easing.

    These can and do disagree, sometimes for months. Reporting that “the yield curve inverted” without specifying which spread is describing a choice of measure as though it were a fact about the world.

    The limitations that get ignored

    The lead time is long and inconsistent. Historically the gap between inversion and recession onset has varied widely — often somewhere between roughly six months and two years. An indicator with a range that wide is close to useless for timing anything. Positioning defensively at the moment of inversion has meant sitting out significant market gains in more than one cycle.

    The sample is small. Developed economies have experienced a limited number of recessions since reliable yield data begins. Claims of near-perfect predictive accuracy rest on a handful of observations — a sample from which strong statistical confidence cannot honestly be drawn.

    False positives exist. There have been inversions not followed by recession within any reasonable window, and the historical record contains judgement calls about what counts as a “real” inversion and how long it must persist.

    The term premium may have changed. Sustained central bank bond purchasing, regulatory demand from banks and insurers for high-quality collateral, and global demand for safe assets have all plausibly compressed term premia. A lower structural term premium means the curve inverts on smaller shifts in rate expectations — which would mechanically raise the false-positive rate without any change in economic conditions.

    Un-inversion is the underrated signal. In several past cycles, recession began not while the curve was inverted but shortly after it steepened back — because steepening reflects the market pricing imminent rate cuts in response to visible deterioration. Treating re-steepening as the all-clear inverts the historical pattern.

    How to use it responsibly

    The yield curve is best understood as one input among several rather than a standalone forecast. Read alongside credit spreads, bank lending surveys, unemployment claims and new orders, it contributes real information: it tells you that a market with capital committed expects policy to ease.

    Read alone, as a binary recession switch, it will mislead — not because the relationship is fake, but because it was never precise enough to carry that weight.

    Related reading

    For why rate expectations matter so much to the real economy, see how central bank rate decisions reach the real economy.

    This article is general information and journalism, not investment advice. See our Editorial Policy.