Wrigital
ConsultancyConsultancyFinPrintFinPrintClient Intelligence EngineThe EngineUnverified Answer CountAnswer CountAboutAboutBlogBlogBook an assessment
Brilliant - But I Could Never Put This In Front Of A Client

Brilliant - But I Could Never Put This In Front Of A Client

By Terry Martin10 min read3 August 2026

Brilliant - But I Could Never Put This In Front Of A Client

On the night of 22 October 1707, four ships of the Royal Navy struck rock off the Isles of Scilly in the dark and went down within sight of the English coast. Something close to two thousand men drowned, including the fleet's own admiral, and the fleet went down for a reason that had nothing to do with weather or seamanship. Nobody aboard could say, with any confidence, how far east or west of home the ships actually were.

Latitude, distance north or south, had been solvable for centuries by anybody who could measure the height of the sun or the pole star. Longitude, distance east or west, had no equivalent trick. A ship's position east or west could only be recovered by knowing two clocks at once: the local time where the ship happened to be, read easily enough from the sun overhead, and the time back home in Greenwich, at the very same moment. Compare the two and the gap between them, converted at fifteen degrees an hour, gave the answer. The problem was never the arithmetic. The problem was carrying Greenwich time across an ocean, in a box, for months, without it drifting.

Parliament answered the wreck with the Longitude Act of 1714, and put up a prize worth a fortune to whoever solved it. The scientific establishment of the day had a settled opinion on where the solution would come from, and the opinion was unanimous. It would come from the sky. The moon moves against the background of stars at a measurable rate, and a sufficiently precise set of tables could turn the moon's position into Greenwich time, wherever you happened to be standing. This was called the lunar distance method, and every serious astronomer in Europe worked on it for decades, because everybody already knew the alternative was hopeless.

The alternative was a clock. And every gentleman who tried it proved the astronomers right.

Pendulum clocks were the accurate clocks of the age, regulated by the steady swing of a weight on a rod, and a pendulum clock brought aboard ship behaved the way pendulum clocks behave when the surface underneath them will not hold still. The roll of the vessel threw the swing off. Changes in temperature between the tropics and the North Atlantic changed the length of the metal rod and the timing along with it. Humidity swelled the wood of the case. Every ship's captain, gentleman inventor and mathematician who carried a fine clock to sea for the next several decades watched it lose or gain minutes within days, and every failed voyage added one more data point to a conclusion nobody thought still needed testing. A clock could not keep time at sea. The Board of Longitude, the body set up to judge the prize, employed astronomers almost exclusively and viewed clockmakers as tinkerers wasting the Board's patience on a mechanism everybody already understood the limits of.

Nobody serious asked a different question, which was whether the thing failing at sea had to be a pendulum clock at all.

John Harrison was a Yorkshire carpenter with no formal training and an obsession with clocks that ran without a pendulum and without the friction a pendulum depends on. Across more than thirty years and four increasingly refined machines, he built a mechanism that kept time using a spring and a balance wheel rather than a swinging weight, compensated the whole assembly against temperature with strips of two metals bonded together, and reduced the friction in the moving parts until the thing barely needed oiling. His fourth machine, finished in 1759, was not a smaller pendulum clock. It was not a clock built to the same design as everything the Board had already dismissed. It was built the way a watch is built, sealed and self-contained, indifferent to the roll of a deck in a way nothing tested before it had been designed to be.

It went to sea in 1761, on a voyage to Jamaica, and came back eighty-one days later having lost five seconds.

Fifty years of failed pendulum clocks had taught the Board of Longitude a lesson that was one category too wide. The lesson they had learned was that a clock could not solve longitude. The lesson that was actually true was that the specific clock everybody kept testing could not, because it had been built for a study and a mantelpiece, not for a deck in a swell. Harrison did not refute fifty years of failed experiments. He built the first instance of a different kind of instrument, and the failed experiments had never had anything to say about that instrument, because it did not yet exist.

The same test, run for free, every evening

Somewhere in most advice firms this year, on an evening or a weekend, a principal has sat down with a made-up client scenario and put it to a large language model, out of curiosity, to see whether the thing everybody is talking about could actually do the job.

The output usually reads well. The reasoning usually sounds sound. And then the question that ends the session arrives, most often in exactly this shape: brilliant, but I could never put this in front of a client. There is no source attached to the recommendation that can be checked. There is no record of what was considered and rejected. The arithmetic, when there is any, cannot be shown to have been calculated rather than guessed at, fluently, in the manner of everything else the model produces. Each of those gaps gets tested privately, on a free tool, outside any system the firm actually runs, and each session ends the same way. The conclusion that gets carried into Monday morning, and often into firm policy, is that artificial intelligence cannot be used in regulated advice work.

That conclusion is a pendulum clock conclusion. It generalises a specific failure into an entire category, and the category is the wrong size.

A general-purpose model of the kind available to anybody with a browser was built to hold a fluent conversation and complete a pattern convincingly, which is a genuinely difficult engineering achievement and not the same achievement as verifying a source, executing an exact calculation, or leaving a record of what it considered and set aside. Those are not settings hidden somewhere in the interface, waiting to be found by a cleverer prompt. They are architecture decisions, made or not made before the product existed at all, and a tool built without them will not develop them under interrogation, however good the interrogation is. Testing a consumer chat tool for an audit trail is closer to testing a family car for airworthiness than it looks. The car failing to fly says nothing whatsoever about flight.

What the evidence says

The behaviour itself is well documented. Survey work on UK financial services this year found three quarters of firms already using some form of artificial intelligence, with more than four in ten of those firms admitting only a partial understanding of the technology they had adopted. Separate research into workplace habits more broadly found a majority of UK employees using consumer AI tools their employer had never sanctioned, most of them weekly, and close to half doing so simply because the same tools were already part of their life outside work. The evening testing session is not an eccentric habit. It is close to the median behaviour.

The architectural point holds up under scrutiny too. A model of this kind is trained to predict the next plausible piece of text, not to execute a verification rule the way a database lookup or a calculator does, which is why it can write fluently about arithmetic while remaining unreliable at performing it, and why it can describe a source persuasively without any mechanism forcing that source to actually exist. International safety research on these systems continues to report false and unreliable output as a standing feature of the category, not a bug awaiting a patch.

Ask the industry what worries it and the answer lines up with the same gap. Recent industry-wide research into financial services found data privacy ranked as a leading concern by close to three quarters of respondents, with unreliable or hallucinated output close behind at roughly seven in ten. The concern is not a fringe view held by the cautious. It is close to consensus, and it is a concern about exactly the properties the evening testing session keeps failing to find.

What the evidence does not offer is a single documented case of a firm's written policy changing overnight because of one bad private test. What the wider literature on how organisations respond to new and hard-to-govern technology does show is a consistent pattern: risk that looks difficult to bound gets managed by restriction, and an informal failure hardens into a formal ban far more often than it gets narrowed into a more precise question. The mechanism fits what advice firms describe, even without a single traceable case proving the causal chain.

What this means for your firm

Separate the conclusion from the category before either one goes into a committee paper. A test that failed was a test of one tool, not of the field, and the difference is worth a sentence in the minutes rather than an assumption nobody revisits.

Write down what your regulated workflow actually requires before you judge anything against it. Name the source verification, the calculation and the record of what was considered and set aside as explicit requirements, the way you would specify requirements for any other system entering the firm, rather than testing a tool informally and treating whatever it happens to lack as proof the requirement was never available.

Ask any AI system you consider, existing or proposed, to show you where each of those three requirements is met by design rather than by promise. A system built to verify a source can show you the retrieval. A system built to calculate can show you the calculation running outside the language model rather than inside its guesswork. A system with nothing to show you at any of the three has told you what it is, and you have not yet tested the category.

Revisit the ban, if your firm has one, on a fixed date rather than never. A restriction adopted in response to one tool's failure is a reasonable holding position and a poor permanent one, and the honest version of the policy names the date it will be looked at again.

The instrument, not the lesson

Nobody at the Board of Longitude was foolish for watching pendulum clocks fail at sea for fifty years. They were foolish only in the size of the lesson they let themselves learn from it. The eighty-one days that mattered were not the eighty-one days at sea in 1761. They were the thirty years before it, spent building a mechanism nobody had specified because nobody yet knew such a mechanism was the missing piece rather than a smaller version of the piece already tried and already found wanting.

Your evening test was not wrong. It tested a tool that was never built to survive the questions you were right to ask it. The tool that can answer them is not a better version of the one you already dismissed. It has not been built the same way at all, and the question worth asking on Monday morning is not whether artificial intelligence belongs in regulated advice, but whether the thing sitting on your desk was ever the instrument the question required.

Ready to see FinPrint or book an assessment?

Start the free countBook a call

Accountable AI for regulated firms.

  • Consultancy
  • FinPrint
  • Client Intelligence Engine
  • Unverified Answer Count
  • About
  • Blog
  • Book a call
PrivacyTerms© Wrigital Ltd