HeadlinesBriefing favicon HeadlinesBriefing.com

AI Chatbots Give Wrong Financial Answers 57% of Time, Study Finds

Financial Times Companies •
×

Research from technology firm Saturn shows that popular AI models from Chat GPT, Claude, Copilot, Grok and Gemini provide incorrect answers to financial queries 57% of the time on average. For complex questions requiring multiple calculations, error rates rise to 88%, with some models failing 99% of the time. The study tested over 100 money-related questions across 18 AI models, totaling more than 10,000 queries.

Errors included calculation mistakes, omitted tax changes, and hallucinated rules. In one case, Claude Haiku 4.5 gave wrong pension tax guidance risking a £17,500 HMRC charge. Another error had Claude inventing a rule allowing student loan repayment stops when moving abroad.

Paid models outperformed free ones, and newer versions were more accurate. The best performer, Claude Opus 5 in reasoning mode, still erred 39% of the time. Sarah Coles, head of personal finance at AJ Bell, said advisers increasingly see clients acting on baffling AI suggestions.

While chatbots can help with budgeting and research, the Financial Conduct Authority found one in five UK adults is open to AI making financial decisions, with younger investors especially trusting. Saturn's Amal Jolly warned that AI financial advice remains unregulated, leaving consumers without protections or compensation available from human advisers.