NextFin

AI Chatbots Give Wrong Answers to Financial Queries 'Most of the Time'

Summarized by NextFin AI
  • AI chatbots fail 57% of financial queries in a comprehensive test of 18 models against 121 questions, with the best system still wrong 39% of the time.
  • Free models underperform paid versions: 63% error rate versus 49%, and on the hardest questions free models failed 93% of the time.
  • 26% of consumers trust general-purpose AI for financial advice, with 56% of 18-40 year-old investors trusting AI tools more than traditional media sources.
  • The FCA is weighing regulation of general-purpose AI tools, as current chatbots dispensing financial guidance carry no regulatory safety net for consumers.

NextFin News - Artificial intelligence chatbots give wrong answers to financial questions most of the time, with mainstream models failing on 57% of queries in one of the most comprehensive tests of AI money advice to date - and the best-performing system still got it wrong 39% of the time. The gap between how much investors trust these tools and how often they are wrong has become the central risk in the race to automate financial guidance.

The Numbers: A Majority Error Rate Across Every Major Model

The report, titled Artificial Authority: Should you trust AI to deliver financial advice?, tested 18 widely used AI models - including versions of ChatGPT, Claude, Copilot, Grok and Gemini - against 121 financial questions covering debt management, mortgages, pensions, tax, savings and student loans. Each question was asked five times to check for consistency, producing more than 10,000 individual responses for analysis. The models delivered accurate answers in only 43% of cases.

On harder questions, the error rate climbed to an average of 88%, with some models answering as many as 99% of the most complex queries incorrectly. Free-to-use models performed worse than their paid counterparts: 63% of free-model responses were wrong, compared with 49% for paid versions. On the hardest questions, free models failed 93% of the time.

Model-by-model, the spread was wide but the direction was uniform. The worst performer, Claude Haiku 4.5, made mistakes in 82% of answers. Google's Gemini 3.1 Pro followed at 73%. xAI's Grok 4.5 was wrong 59% of the time, and ChatGPT 5.6 Luna 58%. Even the strongest system tested, Claude Opus 5 in reasoning mode, still produced incorrect answers in 39% of cases - wrong roughly two out of every five times.

The errors were not abstract. They included calculation mistakes, omitted risk warnings, failure to account for upcoming tax changes, and the invention of financial rules that do not exist. In one case, a pension tax error from Claude Haiku 4.5 could have left a saver facing a £17,500 charge from HM Revenue & Customs. In others, models recommended paying high-interest debt before priority bills such as rent and council tax - advice that could expose households to eviction, bailiff action or legal proceedings. One model told a graduate moving abroad that they could simply stop student loan repayments; another said a mortgage payment holiday would not affect a borrower's credit score. Both statements were wrong.

"The low quality of financial advice from mainstream AI models risks leading to widespread consumer harm," said Amal Jolly, chief executive of Saturn, the financial technology firm behind the research. "Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money."

The research lands as the UK's Financial Conduct Authority weighs whether general-purpose AI tools should fall within its regulatory perimeter. The FCA's Mills Review of AI in retail financial services, published in July 2026, found that around 26% of consumers trust general-purpose tools such as ChatGPT, Claude or Gemini for financial advice, often with limited awareness that formal routes to recourse will not apply. At present, chatbots that happen to dispense financial guidance carry no regulatory safety net; only tools specifically set up to provide financial advice would likely come under the watchdog's remit. Saturn is calling for the FCA to act before consumers face significant losses - a position that carries a commercial dimension, since the firm sells AI infrastructure to regulated advice businesses and stands to gain from stricter rules on consumer-facing chatbots.

Why Probabilistic Models Fail at Deterministic Finance

The core problem is architectural, not a matter of scale. Large language models are probabilistic engines trained to predict the next token in a sequence. Finance, by contrast, is a domain of deterministic rules: a tax threshold is either crossed or it is not; a pension allowance is either exceeded or it is not; a student loan repayment obligation either survives emigration or it does not. When a system that deals in likelihoods is asked a question with a single correct answer, it will produce a plausible-sounding answer even when it has no reliable basis for one. This is not a bug that more training data fixes; it is the native operating mode of the technology.

That mismatch explains why accuracy did not improve smoothly with model sophistication. The best-performing system in the test was a reasoning-mode model - architecture designed to work through problems step by step - yet it still failed more than a third of the time. The takeaway is uncomfortable for the industry's scaling thesis: bigger models get better at sounding authoritative, but authority is not the same as correctness. A confident, well-structured, empathetically worded answer can be wrong in exactly the same way a sloppy one is, and the polish makes the error harder to spot.

The consistency check embedded in the methodology sharpens the point. Each question was asked five times. A rule-based engine returns the same answer five times because it is executing the same rule. A probabilistic model can return five different answers to the same question, which means the user's outcome depends partly on which version of the answer they happened to receive. In finance, where a single decision compounds over decades, that variance is not a curiosity - it is a material source of risk.

There is also a data-horizon problem. Tax rules change, pension allowances shift, and mortgage products expire. A model's knowledge is fixed at its training cutoff unless it is connected to live authoritative sources through retrieval. None of the models in the test was grounded in a live regulatory database; they were answering from static, internal weights. That is why "failure to account for upcoming tax changes" appeared as a recurring error category rather than an edge case. Until AI financial tools are tethered to primary sources - HMRC guidance, FCA rules, pension scheme documents - they will keep hallucinating rules that no longer exist or never existed at all.

The Trust Gap: Why the Error Rate Matters More Than It Should

If chatbots were wrong 57% of the time but everyone knew it, the harm would be contained. The danger is that users do not know, or believe they are protected when they are not. The FCA's Mills Review found that around 26% of consumers trust general-purpose AI tools for financial advice. Among 18- to 40-year-olds who own or are considering investments, a separate FCA survey found 56% trust AI tools - more than TV and radio (47%), the press (46%) or social media influencers (29%). Two-thirds expect to lean on AI even more over the coming year.

The misconceptions run deeper than trust. Almost half (44%) of 18- to 40-year-olds wrongly believe that financial information from AI chatbots is regulated. Nearly a third (32%) think they would be entitled to compensation from the Financial Services Compensation Scheme or the Financial Ombudsman Service if AI-driven advice caused them losses. Thirty-eight per cent consider it acceptable to make an investment decision based solely on AI outputs. These are not marginal beliefs; they describe a cohort that has normalized AI as an authority figure in the one domain where being wrong has a direct, compounding cost.

The mechanism of harm is psychological as much as numerical. The study noted that large language models deliver financial advice with a convincing tone of confidence and care, often wrapped in disclaimers such as "always do your own research." For a financially literate user, that disclaimer is a red flag. For the user who needs the advice most - the one without the time, vocabulary or confidence to cross-check - the disclaimer performs no protective function at all. It is theatre, and it is worse than theatre because it gives the impression of responsibility while assigning none.

This is the second-order risk that the headline error rate understates. The first-order problem is that the answers are wrong. The second-order problem is that the wrong answers arrive packaged as trustworthy, which suppresses the very behavior - verification - that would catch them. A user who expects a 50% error rate will double-check. A user who believes the tool is regulated will not.

The Counter-Thesis: A Starting Point, Not a Destination

The strongest argument against tighter rules is that AI remains useful for financially literate users who treat it as a research assistant rather than an adviser. Independent academic work has found that AI offers highly structured and practical guidance and can be a valuable aid for fact-finding and brainstorming - provided the user already knows enough to spot errors and cross-check data. The FCA itself has said AI can help people research companies, understand jargon and explore options before making a decision, provided they "continue to use your own judgement."

There is also an equity argument. Human financial advice in the UK is expensive and inaccessible to large parts of the population. If AI can close even part of that advice gap at near-zero marginal cost, restricting it could entrench inequality by reserving quality guidance for the wealthy. Saturn's own mission is built on this premise: using technology to make human-led advice accessible to a billion people.

Both points are valid, and neither survives contact with the numbers in this study. A tool that is wrong 57% of the time is not a "starting point" for the median user; it is a coin flip dressed as expertise. The equity argument only works if the AI is grounded in authoritative sources and its error rate is low enough that a non-expert can safely act on it. At 39% wrong even for the best model, we are not close to that threshold. The right policy conclusion is not to ban AI from finance but to require that any AI presenting itself as financial guidance disclose its error rate on standardized tests, cite its sources, and - critically - be tethered to live regulatory data rather than static training weights.

What to Watch: The Signal That Would Change the Thesis

The central judgment here is that AI's financial-advice problem is structural, not cyclical. It will not self-correct through routine model upgrades, because the failure mode is baked into the probabilistic architecture and the absence of source grounding. That call can be falsified. The signal to watch is this: if frontier models equipped with retrieval-augmented grounding in official tax, pension and regulatory databases can drive error rates below 10% on standardized financial benchmarks while maintaining answer consistency across repeated identical queries, the structural thesis is wrong and the problem becomes one of deployment discipline rather than architecture.

By time horizon, the outlook splits. In the short term, expect more studies, more headline error rates, and growing pressure on the FCA to clarify the perimeter - the regulator has already signaled that general-purpose chatbots sit outside it. In the medium term, the winners will be hybrid systems: AI front ends that retrieve from authoritative sources and route complex cases to human advisers, rather than attempting to answer from internal weights alone. In the long term, the question is whether AI becomes a regulated component of the advice chain - subject to accuracy standards, audit trails and compensation obligations - or remains a wildcat channel that regulators are forced to tolerate because it cannot be switched off.

For investors, the exposure is asymmetric. The companies building ungrounded consumer-facing advice chatbots carry regulatory and reputational risk that is not yet priced in. The companies selling compliance, retrieval and audit infrastructure to regulated advice firms stand to benefit from whatever rules emerge. The FCA's next move on the perimeter is the catalyst; the Saturn study is the evidence base that makes inaction harder to defend.

The final takeaway is sharper than the headline suggests. The problem with AI financial advice is not that chatbots are stupid; it is that they are persuasive. A wrong answer you distrust is harmless. A wrong answer you believe - delivered with warmth, structure and a disclaimer you do not read - is how people lose money they cannot afford to lose. Until the technology can tell the difference between sounding right and being right, the burden of verification falls on the user, and the user is precisely the person least equipped to carry it.

Explore more exclusive insights at nextfin.ai.

Insights

Why do AI models fail at finance?

What makes finance rules deterministic?

How do probabilistic AI models work?

Why do chatbots hallucinate money rules?

What was the overall AI error rate?

Which model performed the best?

How often do free models fail?

Do users trust AI money advice?

Are AI chatbots currently regulated?

What did the FCA Mills Review find?

What is Saturn calling on FCA now?

Will hybrid AI systems win out?

Can error rates drop below 10%?

Will AI become regulated advice?

What signals would change the thesis?

Why is answer consistency a major risk?

Why are AI disclaimers ineffective here?

Who bears the verification burden?

Did paid models beat free ones?

How does AI compare to human advice?

Search
NextFinNextFin
NextFin.Al
No Noise, only Signal.
Open App