The work your clients can see is the work AI does best
By the measure BP used to judge itself, the Texas City refinery was in decent shape. The site reported personal injury rates, and personal injury rates were what travelled up to management: slips, cuts, falls, days away from work. Those numbers had been moving in the right direction.
On 23 March 2005, an isomerisation unit at the refinery overfilled during start-up, released hydrocarbon vapour, and ignited. Fifteen people died. More than a hundred and seventy were injured.
The Baker Panel, which reviewed BP's refining safety afterwards, found something more troubling than a failed valve. BP had been treating personal injury rates as a proxy for process safety. The two measure different things. One tells you whether a fitter trips over a hose. The other tells you whether a refinery explodes. What the proxy produced, in the panel's words, was a false sense of confidence.
Nobody in that reporting chain was lying. The instrument worked. It measured something real, reported it, and showed improvement year on year. The trouble was that the thing it failed to measure had no instrument pointed at it at all.
Underneath the reported numbers sat the pattern the panel found wherever a refinery ran short of resource: maintenance deferred, inspections and testing falling behind. That is work which produces nothing visible when it goes well. A pressure vessel inspected and found sound generates no work order, no repair, no line on a summary. The engineer who spends a day confirming that nothing is wrong has, on the record, spent a day producing nothing.
So the maintenance function sat in a bind that any principal who has defended a fee will recognise. The visible half of the job, fixing what broke, produced evidence of itself. The invisible half, judging what did not need fixing and confirming what was sound, produced silence. When money got tight, one half of the work could speak up. The other could not.
What the story is about
Take the refinery out of it and the shape is this. Organisations measure the part of the work that generates observable output. The part that carries the risk generates none. Over time the measured part becomes the definition of the work, and the unmeasured part becomes a cost with no defence.
That is the position advice firms are being moved into, and the thing doing the moving is a chatbot.
The client with three printed pages
Picture the meeting. A client sits down for an annual review with three pages printed from ChatGPT. Two of them are about the recommendation. The third is about the fee. He is not hostile. He has done some homework and he wants to talk it through, which is behaviour any adviser would have welcomed five years ago.
Or the version with no printout. A client says nothing about AI, agrees with everything, and it becomes clear in the third question that she checked the answer before arriving and is confirming a match.
If neither meeting has happened yet, the numbers suggest it will.
Lloyds put the number at 28.8 million UK adults using AI to help manage money in a twelve-month period, with 37% using it for investment research and recommendations. The detail that matters more sits in the City AM coverage of the STRAT7 work: around one in ten turn to AI first. Which means the rest are using it somewhere else in the journey. Before the meeting. After the recommendation. Alongside the fee schedule.
Run that against your own book and the arithmetic is uncomfortable.
The FCA's own consumer research makes the position sharper. Consumers use AI for advice, and most of them validate what it tells them against other sources of information.
Read that twice, because it contains a demotion nobody announced. In the client's mental model, the adviser has become one of the other sources. The chatbot produces the position. The adviser is consulted to check it.
That is arriving in the same year as a harder fee conversation. NextWealth has the average ongoing advice fee at 83 basis points, up from 77, with 45% of clients reporting that their fees went up. Two pressures, one meeting, and a client who has done ninety minutes of preparation for a conversation you have been having for twenty years.
Why the obvious defence is a wasting asset
The instinct in the room is to reach for accuracy. The tools make things up. They hallucinate. They cannot be relied on.
Every word of that is defensible today and less defensible each quarter, and the reason has a specific shape worth understanding.
On personal-finance question sets, older models scored somewhere in the 50 to 60% range. Newer ones reach 80%. Move to work that resembles the actual job and the picture inverts. Vals AI's Finance Agent v2 has a top score of 57.86%, with no model clearing 58% overall and financial modelling topping out at 23%. A 2026 adversarial markets benchmark put frontier models at 28% without tools and 67.4% with them, against a human baseline of 80%.
Both of those things are true at once, and the gap between them is the entire argument.
The improvement is concentrated in single-step, well-posed questions with public answers. That is the category of question a client asks out loud. Is a 0.83% ongoing charge reasonable. Should I hold a global tracker instead. Is drawdown better than an annuity at my age. The failures are concentrated in multi-step reasoning, precise figures, and adversarial source checking. That is the category of work no client has ever watched anyone do.
So the accuracy defence weakens in the room while remaining true outside it. Every quarter it makes the adviser sound more defensive and less current, and the client goes away with the impression that their adviser is arguing with a tool rather than answering a question.
Condition monitoring in industry followed the same curve. Vibration and thermal systems that produced false positives at rates between 15 and 40% on deployment now run at 3 to 8% after tuning. Anyone who dismissed the technology on false alarms in 2020 was correct then and looks foolish now.
What did not improve is instructive. The survey literature is direct about it: a failure mode absent from the training data gets missed or filed as something familiar. Models lose validity when the process, the tooling, the suppliers or the environment shift underneath them. Overlapping faults defeat systems that handle single faults well, because one signal masks another. And the decision about whether to fix a defect now, monitor it, or run the asset to failure needs cost, safety, production and criticality context that no sensor holds.
The machine detects. The machine does not decide. Twenty-six years of improvement has not touched that boundary, and it is where the value went.
The treadmill
Competing on answers concedes the premise. Win the accuracy argument this quarter and it has to be won again next quarter, on ground that shifts toward the machine each time. A firm that positions itself as the more reliable answer-generator has agreed that the answer is the product, and has entered a contest it will lose by degrees, in public.
There is a second problem. The client who arrives with a printout has, in most cases, a reasonable answer. Arguing with it makes the adviser wrong in the moment as well as defensive.
The catalogue nobody sees
Ask what happened between the client's circumstances and the recommendation, and the honest inventory reads:
- The options screened and rejected, with the reason each one failed
- The checks run that came back clean
- The constraints applied: capacity for loss, tax position, liquidity needs, platform limits, a business owner's cash demands
- The reason for this route rather than the obvious one
None of it appears in the output. The client receives a recommendation, which is the one artefact a chatbot can also produce.
This is not peculiar to advice. An ethnography of general practice found that the non-patient-facing work, interpreting results and crafting referrals, was poorly understood and often unrecognised and undervalued, by patients, by policy makers, and by other clinicians. The mechanism is the same everywhere it has been studied. Clients observe outputs. They cannot observe diagnosis, triage, interpretation or the management of uncertainty. Where value sits in errors avoided, it produces no event to point at.
Aviation solved this in 1978
Nowlan and Heap's report for the US Department of Defense reframed maintenance around whether a task was applicable and effective at minimum cost, rather than around how much preventive work got done. The finding that landed hardest was that scheduled overhauls on turbine engines delivered no reliability or economic benefit, while on-condition maintenance reduced cost and improved reliability.
The durable change was in what maintenance reported. The function moved from work orders completed to failures prevented, risk reduced, cost avoided. It stopped selling activity and started publishing judgement.
That answer existed for twenty-seven years before Texas City, in an adjacent industry, in a public document. It did not travel. Professions do not inherit each other's solutions, and the ones that need them most are the ones running a good year.
What this means for your firm
Instrument the rejections. For your next recommendation, write down the three options you screened out and why. Not for the file. For the client. The rejected route is the clearest evidence that a judgement occurred.
Report the checks that found nothing. The FCA's 2025 review of ongoing advice services expects firms to evidence delivery of everything they contracted to deliver. The review found no systemic issue among the largest firms, and it declined to examine the quality or value of the advice itself. That is the next question, and firms that can only show the annual review happened will have a thin answer.
Name the constraints out loud. A chatbot answering the same question has no access to your client's capacity for loss, their tax position, their liquidity needs or their platform. Say which constraint moved the recommendation. It is the part the tool cannot hold by design.
Separate detection from decision. Let the client bring the printout. Treat it as a detection layer that has done some real work, then show where the decision sat. This costs nothing and it removes the argument.
One trade-off, stated flat. Consumer Duty has already generated documentation that nobody reads, and none of this helps if it lands in the file and stops there. The distinction is the audience. Written for the client, this is the product. Written for the regulator and filed, it is cost with no return.
The question
BP's instrument worked. It measured a real thing, and it improved, and the improvement was true. What it could not do was tell anyone what was happening in the part of the operation that had no instrument on it.
Every advice firm has an instrument pointed at the recommendation. Look at what you send a client after a review and you will find the output, the rationale, and the fee. The deliberation that produced all three is measured by the absence of complaints.
Which part of your work is being measured that way, and how long do you think that holds?
