September 20, 2026
Better not ask an AI bot about money: 57% of answers in the test were incorrect
Even the best of 18 models got every third financial question wrong; the worst missed four answers out of five.

One wrong answer about pension taxes would have cost an investor £17 500 in additional assessments. The model that erred was Claude Haiku 4.5.
The mistake was not isolated. Saturn ran 18 models through 121 financial questions five times each, collected more than 10 000 answers, and found 57% were wrong. The questions covered everyday topics: debt, mortgages, pensions, taxes, savings, and student loans.
Already, 26% of users bring questions about their own money to a regular chat such as ChatGPT or Claude (FCA measurement).
Complex questions break everyone. On questions involving calculations, the share of wrong answers rises to 88%, while Gemini 3.5 Flash and Claude Haiku 4.5 miss 99% of such cases.
A paid plan helps little. Free models were wrong in 63% of answers, paid ones in 49%. On complex questions, the error rate for free models reaches 93%.
The spread between models is large. Claude Opus 5 had the best result at 39% errors; Claude Haiku 4.5 had the worst at 82%.
The mistakes look ordinary. Claude advised stopping student-loan payments when moving abroad, while Gemini claimed mortgage payment holidays do not damage a credit score. The calculations used UK rules: HMRC issues the additional assessment in the pension example.
Who measured it. Saturn Fintech released the "Artificial Authority" report on 16 September. The company sells AI to financial advisers and raised a $15 mln Series A backed by Y Combinator in October 2025.
None of the 18 models fell below a 39% error rate in the measurement.
Source
