Wrong 57% of the Time: What the Saturn Study Really Says About AI Financial Advice
Saturn put 18 AI models through 10,000+ money questions and found wrong answers 57% of the time — rising to 88% on complex queries. A PensionBee survey shows 57% of users would act without checking.
Ask a chatbot about your pension, your mortgage, or your taxes, and you will get an answer. Whether you should trust it is a different question — and this week it got a number attached to it.
A study by UK fintech firm Saturn, reported by the Financial Times, found that popular AI models give wrong answers to financial questions 57% of the time. On more complex queries, the error rate climbs to 88%, and some of the weakest models missed hard questions as often as 99% of the time.
The timing matters. Two days after the FT’s coverage landed, a separate PensionBee survey of 1,000 US adults found that 57% of people who ask chatbots for money advice would act on it without independently verifying the answer. Two different studies, the same number, pointing at the same uncomfortable gap: the tools are wrong more often than their users assume, and the users don’t check.
How the study worked
Saturn’s report, titled Artificial Authority: Should you trust AI to deliver financial advice?, tested 18 AI models — spanning the major consumer assistants including ChatGPT, Claude, Copilot, Grok and Gemini — against 121 financial questions covering debt, mortgages, pensions and tax. Each question was repeated five times to test consistency, pushing the total volume past 10,000 individual queries.
An answer was marked as a failure if it contained a factual error, omitted something material, or left out a risk warning that should have been there. That last criterion is worth pausing on: under UK-style regulated advice, missing a suitability warning isn’t a stylistic flaw — it’s the difference between advice and malpractice.
The headline results, in Saturn’s own scoring:
- 57% of all answers were wrong
- 88% error rate on advanced, multi-step questions
- 99% failure for the weakest models on the hardest questions
- Free models erred 63% of the time versus 49% for paid models — better, but still a coin flip weighted toward failure
- On the hardest questions, free models reached a 93% error rate
The league table nobody wants to top
Saturn named names. Claude Haiku 4.5 came last, wrong on 82% of answers. Gemini 3.1 Pro followed at 73%. Grok 4.5 sat at 59% and ChatGPT 5.6 Luna at 58%. The strongest performer was Claude Opus 5 in reasoning mode — still wrong 39% of the time.
Let that sink in: the best model in the study fails on roughly two out of five money questions. If a human adviser posted those numbers, they would be out of a job and possibly in court.
The errors that cost real money
The aggregate percentages are abstract; the specific failures are not. Three examples from the report stand out:
A £17,500 tax trap. One model’s mistake on UK pension tax rules would have exposed a saver to a £17,500 charge from HMRC had they acted on it.
The debt-priority trap. In debt scenarios, models steered users toward clearing the highest-interest balance first — standard-sounding advice that ignores priority arrears like council tax and rent. Following it can end in eviction or bailiffs, because those debts carry legal consequences that a credit card never will.
A fabricated emigration rule. One model invented a rule allowing graduates to suspend student loan repayments by moving abroad. In reality, emigration can raise monthly student loan payments, not pause them. The model didn’t get a nuance wrong — it conjured a rule out of thin air and stated it with full confidence.
There was also a mortgage example flagged across the coverage: a Gemini model wrongly reassured a borrower that taking a mortgage payment holiday would not damage their credit score, when in reality it could — making it harder to secure competitive rates later.
This is the core failure mode of AI financial advice: not hesitation, but confident specificity about rules that don’t exist. A wrong human adviser at least knows when they’re guessing. A language model produces the same fluent, authoritative tone whether it’s right or wrong.
The trust gap, in numbers
The PensionBee survey — fielded July 18–21, 2026 among 1,000 US adults who use chatbots for personal finance — fills in the other half of the picture:
- 57% would act on chatbot guidance without checking it
- 23% say a chatbot has already given them wrong money information; 5% discovered the error only after acting
- 18% would proceed with an investment allocation the bot suggested; 15% would adopt a chatbot-recommended retirement age — decisions that are hard or impossible to reverse
- Despite 53% reporting moderate or extreme privacy concerns, 32% share monthly spending details, 29% share income, and 3% have shared a passport number, driver’s license number or Social Security number
- 66% of Gen Z would let AI act autonomously on their behalf, versus 47% of Baby Boomers
The generational split cuts both ways. Baby Boomers are the least likely to let AI act unchecked — but also the most complacent about accuracy: 82% of Boomer users say a chatbot has never given them wrong or unsuitable information, the highest share of any generation. Given Saturn’s error rates, the most plausible explanation is not that Boomers get better answers. It’s that they don’t recognize the wrong ones.
Read the provenance before the percentages
Saturn designed the methodology, applied the scoring, and paid for the work — and it is explicitly lobbying the UK’s Financial Conduct Authority to bring AI financial advice inside the regulatory perimeter, an outcome Saturn would commercially benefit from. That doesn’t make the findings wrong, but it means the precise figures are the firm’s own, and should be cited as such rather than treated as an independent benchmark. The absence of independent replication is the real weakness of this whole research area — though with 18 named models and a defined question set, replicating it is now straightforward for anyone without a commercial stake.
The regulatory backdrop is not empty either. The FCA’s own Mills Review found 26% of consumers already trust general-purpose AI tools with financial questions, and the regulator has warned that people taking unregulated AI advice get none of the redress that regulated advice carries. More recent FCA research found that 56% of 18-to-40-year-old investors trust AI tools on investing — ahead of broadcast and print media — while nearly half wrongly believe AI financial advice is already regulated. It isn’t.
Why this keeps happening
The failure pattern in Saturn’s data is structural, not incidental. Financial advice is a domain where correctness depends on jurisdiction-specific rules that change on political timetables, personal circumstances the model can’t see, and edge cases where the legally correct answer diverges from the intuitively obvious one. Models trained on general web text absorb the average answer to the average question — but a pension tax question has no “average” saver, and a 57% error rate is what averaging looks like when the tails matter most.
There’s also an asymmetry of tone. Chatbots answer every question with equal fluency, which systematically overstates the reliability of their weakest domain. Users calibrate trust to how confident the answer sounds, and language models sound identically confident at 39% error and 82% error.
The practical takeaway
None of this means AI is useless for money questions. The models’ strength is explaining concepts — what an ISA is, how compound interest works, what a pension annual allowance means. The danger zone is specific, consequential, jurisdiction-bound advice: how much to withdraw, which debt to clear first, whether to take a payment holiday.
The sensible division of labor for now: use chatbots to understand the question, and use regulated humans — or at minimum, official sources — to settle the answer. Saturn’s own CEO Amal Jolly put it bluntly: “Millions of people are trusting the AI models for money advice, but they are getting wrong answers that can lose them money.”
When the best model in the study is still wrong 39% of the time, “trust but verify” isn’t strong enough. Just verify.
Sources
- [1] https://www.ft.com/content/c0cd359d-df84-4208-a789-ffa864b43666
- [2] https://professionalparaplanner.co.uk/millions-of-britons-at-risk-of-financial-harm-from-flawed-ai-answers/
- [3] https://www.resultsense.com/news/2026-09-14-ai-financial-advice-error-rate/
- [4] https://www.stocktitan.net/news/PBNYF/pension-bee-study-suggests-alarming-ai-personal-finance-nbsqfg2kc3zl.html
- [5] https://www.pensionbee.com/us/research-and-insights/2026-ai-and-your-money-report