How the study measured AI overreliance
Most work on AI overreliance asks whether people accept wrong answers. This paper asks a narrower question: does the option of not answering survive contact with an AI assistant? Declining to answer is a normal move when someone's knowledge runs out. The experiments treat that move as the thing being measured.
The paper at a glance
Study overview
Suspension of judgment, experiment by experiment
Participants answered six hard film questions and could always decline. The table compares the share who declined when AI advice was reachable against the share when it was not. Preregistration here means the hypothesis and analysis plan were published before the data was collected.
| Experiment | Participants | No AI | AI available |
|---|---|---|---|
| Study 1a | 314 | 0.36 | 0.06 |
| Study 1b | 310 | 0.44 | 0.03 |
| Study 2 (no stakes) | 812 | 0.17 | 0.01 |
| Study 3 (no stakes) | 843 | 0.32 | 0.02 |
| Study 4 (no stakes, automatic) | 853 | 0.35 | 0.01 |
The widest gap was Study 1b, where the rate fell from 0.44 to 0.03. That experiment pre-wrote three wrong answers and displayed them, which removes flaky tool behaviour as an explanation. The result still replicated Study 1a.
"In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond." (Abstract) / "Participants in the baseline condition were substantially more likely to suspend judgment than participants who could seek AI advice (0.36 vs. 0.06; one-tailed t-test: t = 8.295, p < .001; see Figure 1a)" (Study 1a) / "Participants in the baseline condition were again substantially more likely to suspend judgment than participants who could seek AI advice (0.44 vs. 0.03; t = 11.772, p < .001; see Figure 1b)" (Study 1b) — from the arXiv paper
Why rational delegation cannot explain the drop
The obvious reading of any deference result is that people handed the work to something better at it. This design closes that door before the experiment starts.
The questions were engineered so the advice was wrong
The authors state plainly that they engineered the questions so the AI advice would be wrong, separating AI use from its accuracy. Whatever made participants stop declining, it was not the advice being good.
That leaves one reading. A plausible-looking answer sitting in front of you changes how you handle your own uncertainty. The paper frames this as a shift in the metacognitive threshold, the point at which a person decides they know enough to speak.
"We engineered the questions so that AI advice was wrong, separating AI use from its accuracy." / "As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer." (both from the Abstract) — from the arXiv paper
Accuracy fell to a third while confidence more than doubled
Pooling the conditions without money at stake, correctness was 27.5% without AI and 9.2% with it. Mean confidence ran the opposite way, 29.6 against 75.9. People answered more often, were right less often, and felt considerably better about it.
This pairing is the part that matters at work. A wrong answer delivered tentatively gets checked. A wrong answer delivered with conviction gets shipped.
"in the absence of monetary incentives, participants answered more questions but were correct about a third as often as when AI was unavailable (pooled correctness: 27.5% vs. 9.2%), while confidence was roughly two and a half times as high (mean confidence: 29.6 vs. 75.9). Without incentives, AI access made people far more assured and far less accurate." (Discussion) — from the arXiv paper
What money changed about AI dependence
Studies 2 through 4 added stakes: $0.10 for each correct answer, $0.10 lost for each wrong one, and nothing either way for declining. The authors had preregistered a prediction that stakes would restore suspension, and that AI would blunt the effect. The data did not agree.
Stakes lifted suspension a little, never back to baseline
Money raised the rate of declining, but nowhere near the no-AI level. In Study 3 the AI condition moved from 0.02 to 0.08, while the no-AI condition moved from 0.32 to 0.41.
The preregistered AI-by-stakes interaction failed to reach significance in Studies 2, 3 and 4 alike. The two effects did not cancel each other out; they acted separately.
"Across Studies 2–4, then, our pre-registered prediction of a negative AI × stakes interaction on suspension was not supported (p = .274, .506, and .784, respectively). AI availability and stakes acted largely independently: AI sharply reduced suspension regardless of stakes, while stakes modestly increased suspension—most clearly in Study 3—regardless of AI." (Study 4) — from the arXiv paper
Correctness improved only where AI was available
Stakes lifted accuracy in the AI conditions and left the no-AI conditions alone. Study 3 went from 0.11 to 0.16 with AI, while the no-AI figures sat at 0.28 and 0.27. Study 4 repeated the pattern, 0.07 to 0.14 with AI against a flat 0.27 without.
The mechanism visible in the data is a drop in how often people asked. Requests fell from 5.44 to 4.93 out of six questions in Study 2, and from 5.27 to 4.53 in Study 3. Money did not sharpen anyone's judgment. It reduced how often they leaned on the model.
"Like Study 2, when AI advice was available, monetary stakes increased correctness (0.11 vs. 0.16; b = 0.055, SE = 0.017, t = 3.164, p = .002). By contrast, when AI advice was unavailable, stakes had no effect on correctness (0.28 vs. 0.27; b = -0.009, SE = 0.022, t = -0.382, p = .703; see Figure 3b)" (Study 3) / "Participants sought AI advice substantially less frequently when stakes were present than when they were absent (4.53 vs. 5.27 times out of six questions; b = -0.744, SE = 0.182, t = -4.083, p < .001; see Figure 3c)" (Study 3) — from the arXiv paper
Advice nobody asked for had the same effect
Real products do not wait to be asked. Search results open with a summary, editors suggest the next sentence, dashboards surface a recommendation. Study 4 was built to match that.
Study 4 took the choice away
Study 4 (N = 853) removed participants' control over whether advice appeared, showing it automatically in the AI conditions. The results tracked Study 3: suspension fell from 0.35 to 0.01 without stakes, and from 0.39 to 0.07 with them.
Nobody has to reach for the assistant for the effect to land. Seeing the answer is enough. That maps directly onto interfaces where an AI summary is the default rather than a feature you opt into.
"Study 4 (N = 853) therefore removed participants' control over whether to receive AI advice. It was identical to Study 3 except that, in the conditions where AI advice was available, it appeared automatically rather than on request." / "Access to AI advice continued to dramatically reduce judgment suspension, both when monetary stakes were absent (0.35 vs 0.01; b = -0.333, SE = 0.029, t = -11.56, p < 0.001) and when they were present (0.39 vs 0.07; b = -0.322, SE = 0.032, t = -10.18, p < 0.001)." (both from Study 4) — from the arXiv paper
Willingness, not capacity
The authors put weight on the fact that money recovered part of the effect. If the behaviour were a hard cognitive limit, incentives would not touch it. People can still reach for their own judgment when something depends on it, which makes this a question of willingness rather than ability.
The paper links the pattern to epistemia, a term for accepting AI output because it reads smoothly and hangs together grammatically, rather than because it was checked. A language model has to emit something for every prompt and never declines on its own. The suggestion is that users inherit that posture. Questions about who gets credit for AI-assisted writing are covered in how AI bylines are being misused.
"the near-elimination of suspension under AI access is not a fixed cognitive limitation but a contextual response that monetary consequences partly correct: people can recruit their own judgment when motivated, so what AI changes is willingness, not capacity." / "These findings can be read through the notion of epistemia … : the tendency to accept AI outputs for their surface plausibility and syntactic coherence rather than through verification." (both from the Discussion) — from the arXiv paper
There is no HTML version of this paper on arXiv, and the per-experiment figures live only in the PDF. Converting it to markdown with the headings and tables intact keeps each number attached to the study and condition it came from.
The finding is not about how accurate AI is. It is about the threshold on the human side. Put a plausible answer in front of someone and they stop saying "I don't know", and their confidence climbs while their hit rate falls. Wrong and certain is the worst combination to have running through a team. The useful response is not a rule about when AI may be used, but keeping an explicit place in the process where declining to answer is still an available move. That the effect partly reversed under stakes is the evidence that such a place would be used.



