Julien Gruhier, director of AI, assesses the benchmarking of Croner Intelligence against standard AI tools, highlighting the importance of trust and accuracy versus the amount of irrelevant information in generic AI responses
Every AI tool a firm can buy now writes well. Fluency was the hard part five years ago; today it is standard. What separates one tool from another is narrower and less glamorous: whether a tax adviser or accountant can act on the output without re-checking every line, verifying every detail.
Croner’s own benchmarking of Croner Intelligence across UK tax queries illustrates the gap. ChatGPT scored 0.99 for coherence and logical flow and 0.96 for clarity – about as well-written as the scale allows. On relevance and completeness, it scored 0.67. The writing was immaculate, but roughly a third of what the accountant actually needed was missing or irrelevant. Nothing in the response tells you which third was wrong.
The error you cannot see
That invisibility, rather than the error rate itself, is the problem. If a tool flagged the portion of an answer it had inferred rather than sourced, the fix would be procedural: verify the flagged parts, rely on the rest.
But ungrounded material is delivered in exactly the same confident tone as material drawn from statute. The only way to verfity it is to check all of it – at which point the time saved has gone.
The stakes go beyond time. Guidance that leaves the office carries the firm’s name and its indemnity. An answer that cannot be traced back to an authority cannot be defended to a client, a regulator or a tribunal, however plausible it looked on screen.
Two design choices, not a smarter model
The first design choice is grounding. Every statement in Croner Intelligence traces to an identifiable source – legislation, case law, HMRC guidance – retrieved before the answer is drafted rather than recalled from training data. Most of the gain comes from that first step, not from a cleverer model.
In the same testing, a system twice as good at finding the right source material cut answers containing something wrong from roughly one in three to one in seven.
The second is a willingness to stop. Where the sources do not settle a point, Croner Intelligence names the facts that are missing instead of choosing the likeliest reading. Generic tools are optimised to sound confident, which is precisely the wrong optimisation when the law permits more than one answer.
Neither choice makes an answer more impressive. Both make it checkable in seconds rather than reconstructable from scratch.
Professional trust rests on knowing what is accurate. Fluency has become a commodity but knowing what is a reliable source is never optional.
*** Find out more about Croner Intelligence for yourself. Book a demo here ***