FT: 18 AI Chatbots Get Financial Answers Wrong 57% of the Time, Fail 88% on Complex Queries
Summary
A Saturn-run study of 18 AI models across ChatGPT, Claude, Copilot, Gemini and Grok, covered by the FT, finds the chatbots give inaccurate financial advice 57% of the time on 121 questions covering debt, mortgages, pensions and tax — with failure rates climbing to 88% on the most complex multi-step queries. A parallel PensionBee survey found 57% of US chatbot-using adults would act on the advice without independent verification.
Shared on Bluesky by 2 AI experts
-
Using AI chatbots to address financial questions could risk large losses, with models from a range of providers giving wrong answers most of the time. The most popular AI models from ChatGPT, Claude, Copilot, Grok and Ge…
View on Bluesky →
Originally reported by ft.com
Read the original article →Original headline: FT: 18 AI Chatbots Get Financial Answers Wrong 57% of the Time, Fail 88% on Complex Queries