ft.com via Hacker News

FT: 18 AI Chatbots Get Financial Answers Wrong 57% of the Time, Fail 88% on Complex Queries

Summary

A Saturn-run study of 18 AI models across ChatGPT, Claude, Copilot, Gemini and Grok, covered by the FT, finds the chatbots give inaccurate financial advice 57% of the time on 121 questions covering debt, mortgages, pensions and tax — with failure rates climbing to 88% on the most complex multi-step queries. A parallel PensionBee survey found 57% of US chatbot-using adults would act on the advice without independent verification.

Shared on Bluesky by 2 AI experts

  • Hypervisible @hypervisible.blacksky.app: 🫡 →
  • Brent Toderian @brenttoderian.bsky.social amplified

    @robertscotthorton.bsky.social

    Using AI chatbots to address financial questions could risk large losses, with models from a range of providers giving wrong answers most of the time. The most popular AI models from ChatGPT, Claude, Copilot, Grok and Ge…

    View on Bluesky →