Julien Gruhier, head of AI at Croner Intelligence, explains how to measure the quality of AI answers from accuracy to compliance, and relevance to sample testing
Most firms assess an AI tool on a handful of sample questions and an impression of whether the answers feel right. That holds up until somebody asks how the tool was evaluated before it went near client work. ‘Is it any good?’ is not one question. It is many and to be useful, each question should only have a yes or no answer.
1. Accuracy and compliance
Accuracy and compliance are critical. Is the right answer in the response at all? Is everything in the response true? Do the true statements avoid misleading the user? An answer can pass any one of those and fail the next.
It is essential to evaluate for consistency with tax rules and legislation, the correct figures, dates, and technical terms, the precision of the response, and the accuracy of the citations of guidance, legislation and case law.
2. Relevance and completeness
In terms of relevance and completeness, it is important to watch out for replies including references to unrelated topics, superfluous comment, misunderstanding of the initial question, and even occasions when the AI tool answers a totally different question. Clarity and readability are also key, for example is the answer easy to understand, is it logical, is it based on fact, not assumptions?
3. Factual claims and reference material
At Croner Intelligence, we extract every distinct factual claim from the answer, compare these against the reference material the AI tool retrieved, and divide the number substantiated by the total. The result is a direct measure of how much of an answer is based on facts, and how much was supplied by the model.
In line with best practice on AI ranking metrics, Recall@K measures the proportion of correctly identified relevant items, in effect whether the right source was retrieved at all.
An AI tool writing fluently about the wrong document has a retrieval problem, and no improvement to the model will fix it.
4. Testing
Testing is a key stage in the rollout of every new feature. And testing has to be fair, it has to be conducted on a level playing field. So the same question is put to every system iteration under the same conditions and is marked against the same checks. This is effectively the foundational question, and it is re-run at each release.
This level of objective measurement and review means Croner Intelligence is an AI tool accountancy firms and tax advisers can use with confidence, and defend to their clients and boards when they ask how the system was assessed.
*** Find out more about Croner Intelligence, the game changing AI tool for accountants and tax professionals. Book a demo here ***