September 21, 2026
AI models get money questions wrong more often than right: 57% of answers are incorrect
UK fintech Saturn ran 18 models through 121 money questions and analyzed more than 10 000 answers.

A wrong answer about pension taxes could leave someone with a 17 500-pound bill from the UK tax authority.
Saturn assembled such scenarios in its report, “Artificial Authority” (September 2026). The questions cover everyday issues, from debt and mortgages to pensions and student loans. Each was put to a model five times.
On average, the models were wrong in 57% of answers. On difficult questions, the average error rate rises to 88%, while only individual models reach 99%.
Claude Opus 5 in reasoning mode delivered the best result, with 39% incorrect answers. Claude Haiku 4.5 performed worst, at 82%.
ChatGPT 5.6 Luna was wrong in 58% of answers, Grok 4.5 in 59%, and Gemini 3.1 Pro in 73%.
A paid plan helps. Free models were wrong in 63% of answers, paid ones in 49%. On the hardest questions, free models missed the mark 93% of the time.
The mistakes repeat: arithmetic, missed risk warnings, outdated tax rules, and references to regulations that do not exist. In debt scenarios, models advised paying off the most expensive loans first rather than rent and council tax, putting people on the path to eviction and enforcement officers.
Who did the counting. Saturn is not an independent lab but a UK fintech founded in 2023. The company sells an AI platform to financial advisers and has raised $15 million.
Saturn provides its full report only through a form on its website.
Source
