sakutto
Generative AI

AI Overreliance Cut Accuracy From 27.5% to 9.2%

AI LiteracyResearchCognitive Bias
AI Overreliance Cut Accuracy From 27.5% to 9.2%

How the study measured AI overreliance

Most work on AI overreliance asks whether people accept wrong answers. This paper asks a narrower question: does the option of not answering survive contact with an AI assistant? Declining to answer is a normal move when someone's knowledge runs out. The experiments treat that move as the thing being measured.

The paper at a glance

Study overview

Title
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
Authors
Chiara Marcoccia / Walter Quattrociocchi / Valerio Capraro
Posted
15 July 2026 (arXiv:2607.13562)
Scale
5 experiments, N = 3,132 (4 preregistered, 1 direct replication)

Suspension of judgment, experiment by experiment

Participants answered six hard film questions and could always decline. The table compares the share who declined when AI advice was reachable against the share when it was not. Preregistration here means the hypothesis and analysis plan were published before the data was collected.

ExperimentParticipantsNo AIAI available
Study 1a3140.360.06
Study 1b3100.440.03
Study 2 (no stakes)8120.170.01
Study 3 (no stakes)8430.320.02
Study 4 (no stakes, automatic)8530.350.01

The widest gap was Study 1b, where the rate fell from 0.44 to 0.03. That experiment pre-wrote three wrong answers and displayed them, which removes flaky tool behaviour as an explanation. The result still replicated Study 1a.

View official source →
"In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond." (Abstract) / "Participants in the baseline condition were substantially more likely to suspend judgment than participants who could seek AI advice (0.36 vs. 0.06; one-tailed t-test: t = 8.295, p < .001; see Figure 1a)" (Study 1a) / "Participants in the baseline condition were again substantially more likely to suspend judgment than participants who could seek AI advice (0.44 vs. 0.03; t = 11.772, p < .001; see Figure 1b)" (Study 1b) — from the arXiv paper

Why rational delegation cannot explain the drop

The obvious reading of any deference result is that people handed the work to something better at it. This design closes that door before the experiment starts.

The questions were engineered so the advice was wrong

The authors state plainly that they engineered the questions so the AI advice would be wrong, separating AI use from its accuracy. Whatever made participants stop declining, it was not the advice being good.

That leaves one reading. A plausible-looking answer sitting in front of you changes how you handle your own uncertainty. The paper frames this as a shift in the metacognitive threshold, the point at which a person decides they know enough to speak.

View official source →
"We engineered the questions so that AI advice was wrong, separating AI use from its accuracy." / "As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer." (both from the Abstract) — from the arXiv paper

Accuracy fell to a third while confidence more than doubled

Pooling the conditions without money at stake, correctness was 27.5% without AI and 9.2% with it. Mean confidence ran the opposite way, 29.6 against 75.9. People answered more often, were right less often, and felt considerably better about it.

This pairing is the part that matters at work. A wrong answer delivered tentatively gets checked. A wrong answer delivered with conviction gets shipped.

View official source →
"in the absence of monetary incentives, participants answered more questions but were correct about a third as often as when AI was unavailable (pooled correctness: 27.5% vs. 9.2%), while confidence was roughly two and a half times as high (mean confidence: 29.6 vs. 75.9). Without incentives, AI access made people far more assured and far less accurate." (Discussion) — from the arXiv paper

What money changed about AI dependence

Studies 2 through 4 added stakes: $0.10 for each correct answer, $0.10 lost for each wrong one, and nothing either way for declining. The authors had preregistered a prediction that stakes would restore suspension, and that AI would blunt the effect. The data did not agree.

Stakes lifted suspension a little, never back to baseline

Money raised the rate of declining, but nowhere near the no-AI level. In Study 3 the AI condition moved from 0.02 to 0.08, while the no-AI condition moved from 0.32 to 0.41.

The preregistered AI-by-stakes interaction failed to reach significance in Studies 2, 3 and 4 alike. The two effects did not cancel each other out; they acted separately.

View official source →
"Across Studies 2–4, then, our pre-registered prediction of a negative AI × stakes interaction on suspension was not supported (p = .274, .506, and .784, respectively). AI availability and stakes acted largely independently: AI sharply reduced suspension regardless of stakes, while stakes modestly increased suspension—most clearly in Study 3—regardless of AI." (Study 4) — from the arXiv paper

Correctness improved only where AI was available

Stakes lifted accuracy in the AI conditions and left the no-AI conditions alone. Study 3 went from 0.11 to 0.16 with AI, while the no-AI figures sat at 0.28 and 0.27. Study 4 repeated the pattern, 0.07 to 0.14 with AI against a flat 0.27 without.

The mechanism visible in the data is a drop in how often people asked. Requests fell from 5.44 to 4.93 out of six questions in Study 2, and from 5.27 to 4.53 in Study 3. Money did not sharpen anyone's judgment. It reduced how often they leaned on the model.

View official source →
"Like Study 2, when AI advice was available, monetary stakes increased correctness (0.11 vs. 0.16; b = 0.055, SE = 0.017, t = 3.164, p = .002). By contrast, when AI advice was unavailable, stakes had no effect on correctness (0.28 vs. 0.27; b = -0.009, SE = 0.022, t = -0.382, p = .703; see Figure 3b)" (Study 3) / "Participants sought AI advice substantially less frequently when stakes were present than when they were absent (4.53 vs. 5.27 times out of six questions; b = -0.744, SE = 0.182, t = -4.083, p < .001; see Figure 3c)" (Study 3) — from the arXiv paper

Advice nobody asked for had the same effect

Real products do not wait to be asked. Search results open with a summary, editors suggest the next sentence, dashboards surface a recommendation. Study 4 was built to match that.

Study 4 took the choice away

Study 4 (N = 853) removed participants' control over whether advice appeared, showing it automatically in the AI conditions. The results tracked Study 3: suspension fell from 0.35 to 0.01 without stakes, and from 0.39 to 0.07 with them.

Nobody has to reach for the assistant for the effect to land. Seeing the answer is enough. That maps directly onto interfaces where an AI summary is the default rather than a feature you opt into.

View official source →
"Study 4 (N = 853) therefore removed participants' control over whether to receive AI advice. It was identical to Study 3 except that, in the conditions where AI advice was available, it appeared automatically rather than on request." / "Access to AI advice continued to dramatically reduce judgment suspension, both when monetary stakes were absent (0.35 vs 0.01; b = -0.333, SE = 0.029, t = -11.56, p < 0.001) and when they were present (0.39 vs 0.07; b = -0.322, SE = 0.032, t = -10.18, p < 0.001)." (both from Study 4) — from the arXiv paper

Willingness, not capacity

The authors put weight on the fact that money recovered part of the effect. If the behaviour were a hard cognitive limit, incentives would not touch it. People can still reach for their own judgment when something depends on it, which makes this a question of willingness rather than ability.

The paper links the pattern to epistemia, a term for accepting AI output because it reads smoothly and hangs together grammatically, rather than because it was checked. A language model has to emit something for every prompt and never declines on its own. The suggestion is that users inherit that posture. Questions about who gets credit for AI-assisted writing are covered in how AI bylines are being misused.

View official source →
"the near-elimination of suspension under AI access is not a fixed cognitive limitation but a contextual response that monetary consequences partly correct: people can recruit their own judgment when motivated, so what AI changes is willingness, not capacity." / "These findings can be read through the notion of epistemia … : the tendency to accept AI outputs for their surface plausibility and syntactic coherence rather than through verification." (both from the Discussion) — from the arXiv paper

There is no HTML version of this paper on arXiv, and the per-experiment figures live only in the PDF. Converting it to markdown with the headings and tables intact keeps each number attached to the study and condition it came from.

Free ToolPDF to Markdown ConverterConvert PDF content to Markdown format. Auto-detects headings, tables, and lists — ideal for RAG and AI workflows.Try it now →

The finding is not about how accurate AI is. It is about the threshold on the human side. Put a plausible answer in front of someone and they stop saying "I don't know", and their confidence climbs while their hit rate falls. Wrong and certain is the worst combination to have running through a team. The useful response is not a rule about when AI may be used, but keeping an explicit place in the process where declining to answer is still an available move. That the effect partly reversed under stakes is the evidence that such a place would be used.

FAQ

Q. How large was the study?
Five experiments with 3,132 participants in total. Four were preregistered and one was a direct replication. Participants answered six difficult film questions and could decline to answer at any point.
arXiv:2607.13562
In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. arXiv:2607.13562
Q. Would the result disappear if the AI gave good advice?
That is not what the design tests. The questions were built so that the AI advice would be wrong, which separates the effect of using AI from the effect of the advice being accurate. So the change cannot be read as sensible delegation to a reliable tool.
arXiv:2607.13562
We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. arXiv:2607.13562
Q. What happened to accuracy and confidence?
Without monetary incentives, correctness was 27.5% with no AI against 9.2% with AI, roughly a third as often. Mean confidence moved the other way, from 29.6 to 75.9.
arXiv:2607.13562
participants answered more questions but were correct about a third as often as when AI was unavailable (pooled correctness: 27.5% vs. 9.2%), while confidence was roughly two and a half times as high (mean confidence: 29.6 vs. 75.9). arXiv:2607.13562
Q. Does paying people for correct answers fix it?
Partly. With money at stake, participants asked for AI advice less often, answered more accurately, and declined to answer more often. None of it came close to the level seen when AI was simply unavailable.
arXiv:2607.13562
Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. arXiv:2607.13562

Related Tools

Related Tool Categories

Articles