osf.io web signal

AI Advice Cuts People's 'I Don't Know' Answers from 44% to 3%

TL;DR

  • Access to AI advice collapsed participants' willingness to say 'I don't know' from 44% to 3%, while accuracy fell from 27% to 9%.
  • Across five experiments and 3,132 participants, confidence rose from 30% to 76% even as correct-answer rates dropped.
  • Paying people to be accurate barely helped: abstention only rose from 3% to 8%, and accuracy from 9% to 16%.

A new preprint has been making the rounds because it puts a very concrete number on something a lot of people already suspected. When participants were given AI help on hard trivia questions, their willingness to say 'I don't know' collapsed from 44% to 3%, according to reporting in The Register on the PsyArXiv paper by Chiara Marcoccia, Walter Quattrociocchi and Valerio Capraro. Accuracy dropped from 27% to 9% in the same comparison. Confidence went the other way, from 30% to 76%.

The design is what makes those numbers interesting. The researchers deliberately picked obscure visual details from films, the colour of a team's uniform in Bend It Like Beckham was one, and paired participants with a model they knew was usually wrong on that class of question. If the AI had been reliable, the drop in 'I don't know' answers could be waved off as sensible delegation. Here it cannot. Across five experiments and 3,132 participants, per The Next Web's write-up, people took confidently worded bad advice and passed it on with more confidence than they had answering unaided.

The most useful finding for anyone shipping AI features is what happened when the researchers added money. Paying people to be correct did move behaviour, but not much. Willingness to say 'I don't know' climbed from 3% to 8%, and accuracy from 9% to 16%. Both still well under the no-AI baseline of 44% and 27%. That is the part product teams should sit with. If a direct financial penalty for being wrong barely restored the abstention instinct, a soft UI nudge probably will not either.

The honest caveats are worth stating. This is a preprint, the task is trivia rather than a real work decision, and the model was chosen precisely because it fails on this material. What the reporting does not give you is whether the same collapse in abstention holds with a strong, reliable model, or whether repeated exposure eventually recalibrates users. Take the specifics as reported, not settled.

The forward-looking piece is that 'willingness to abstain' is now a measurable, publishable property of a human-plus-AI system, and one that responded weakly even to direct payment. Vendors who invest in calibration UX, forced uncertainty prompts, visible confidence, an explicit unsure button, now have something concrete to point at when they argue that accuracy alone is the wrong metric.

Shared on Bluesky by 2 AI experts