Trust is becoming the most valuable feature in AI

Image

Julien Gruhier, director of AI, assesses the benchmarking of Croner Intelligence against standard AI tools, highlighting the importance of trust and accuracy versus the amount of irrelevant information in generic AI responses

Every AI tool a firm can buy now writes well. Fluency was the hard part five years ago; today it is standard. What separates one tool from another is narrower and less glamorous: whether a tax adviser or accountant can act on the output without re-checking every line, verifying every detail.

Croner’s own benchmarking of Croner Intelligence across UK tax queries illustrates the gap. ChatGPT scored 0.99 for coherence and logical flow and 0.96 for clarity – about as well-written as the scale allows. On relevance and completeness, it scored 0.67. The writing was immaculate, but roughly a third of what the accountant actually needed was missing or irrelevant. Nothing in the response tells you which third was wrong.

The error you cannot see

That invisibility, rather than the error rate itself, is the problem. If a tool flagged the portion of an answer it had inferred rather than sourced, the fix would be procedural: verify the flagged parts, rely on the rest.

But ungrounded material is delivered in exactly the same confident tone as material drawn from statute. The only way to verfity it is to check all of it – at which point the time saved has gone.

The stakes go beyond time. Guidance that leaves the office carries the firm’s name and its indemnity. An answer that cannot be traced back to an authority cannot be defended to a client, a regulator or a tribunal, however plausible it looked on screen.

Two design choices, not a smarter model

The first design choice is grounding. Every statement in Croner Intelligence traces to an identifiable source – legislation, case law, HMRC guidance – retrieved before the answer is drafted rather than recalled from training data. Most of the gain comes from that first step, not from a cleverer model.

In the same testing, a system twice as good at finding the right source material cut answers containing something wrong from roughly one in three to one in seven.

The second is a willingness to stop. Where the sources do not settle a point, Croner Intelligence names the facts that are missing instead of choosing the likeliest reading. Generic tools are optimised to sound confident, which is precisely the wrong optimisation when the law permits more than one answer.

Neither choice makes an answer more impressive. Both make it checkable in seconds rather than reconstructable from scratch.

Professional trust rests on knowing what is accurate. Fluency has become a commodity but knowing what is a reliable source is never optional.

*** Find out more about Croner Intelligence for yourself. Book a demo here ***

Julien Gruhier | Director of search and generative AI, Croner

Julien Gruhier is director of search and generative AI at Croner, and is the driving force behind ...

View profile and articles

0
Be the first to vote

Rate this article

Related Articles
Subscribe