Research Trends 2026-10-11

Ten Minutes of AI Help Made People Worse and Quicker to Give Up Once It Was Taken Away -- a Peer-Reviewed Study Presented at COLM

Three randomized controlled trials with 1,222 participants found that people given ChatGPT while solving fraction or reading problems did better at first, then solved fewer problems and skipped more once the AI was withdrawn. The peer-reviewed paper was presented this week at the Conference on Language Modeling, and co-author Brian Christian argues the effect reaches well beyond classrooms.

Berkeley News reported on October 9, 2026 that the paper "AI Assistance Reduces Persistence and Hurts Independent Performance" had been presented in its updated, peer-reviewed form at the Conference on Language Modeling (COLM). The authors are Grace Liu, Brian Christian, Tsvetomira Dumbalska, Michiel A. Bakker and Rachit Dubey; the team spans UC Berkeley, Carnegie Mellon, MIT, Oxford and UCLA. The arXiv version was last revised on October 3. A preliminary draft first circulated in April and drew wide coverage.

The experiments

  • Fractions, 354 participants. One group solved 15 basic fraction problems alone; the other had ChatGPT open alongside and could even ask it for the answer. The AI group started out more accurate. After 12 problems the tool was removed, and, as Berkeley News puts it, "Almost immediately, the people in that group stopped solving questions accurately."
  • Replication, 667 participants. The AI users "got answers wrong or gave up entirely when the AI assistance was withdrawn," while the group without AI kept going and did better by the end.
  • Reading comprehension, 201 participants. On an SAT reading prompt, persistence and accuracy again dropped when the tool was removed.

In the fractions trial, the rate at which participants skipped questions also rose sharply once assistance ended. The abstract states the core claim: "although AI assistance improves performance in the short-term, people perform significantly worse without AI and are more likely to give up," and "these effects emerge after only brief interactions with AI (approximately 10 minutes)."

The proposed mechanism

The authors argue that today's assistants are "fundamentally short-sighted collaborators - optimized for providing instant and complete responses, without ever saying no." Their explanation is conditioning: "AI conditions people to expect immediate answers, thereby denying them the experience of working through challenges on their own." That matters because, they note, persistence "is one of the strongest predictors of long-term learning."

What Christian draws from it

Christian, a research fellow at Berkeley's Center for Human-Compatible AI and author of The Alignment Problem, said: "This is not a story about fractions and SAT problems." He went on: "This is something that is happening to human knowledge and human expertise from top to bottom and from grade-school students all the way to the leading experts in the world." His concern for science is that a team of ten developing the ideas behind a breakthrough could shrink to "two or three researchers and a chatbot," losing the immersion in data that makes a good scientist. His proposed fix is product design: instead of defaulting to quick answers, assistants could be more instructive, like a tutor, because "We need to work toward a future in which human abilities are, as much as possible, augmented rather than supplanted."

What remains uncertain. The tasks were short and the participants were recruited online; the study measures what happens right after the tool is removed, not whether the effect persists or reverses over weeks. It also tested a general chat assistant, not tutoring modes designed to withhold answers.

Analysis: this is the human side of the agent-reliability stories in this digest. Organisations rolling out assistants usually measure task speed and output quality with the tool on. This study suggests also measuring performance with it off – for onboarding juniors, for incident response when the tool is down, and for any role where people must check the model's work. Training programmes can preserve that capacity by deliberately scheduling unassisted practice, much as pilots still fly manually despite autopilot.

In three randomized trials with 1,222 people, about ten minutes of ChatGPT help improved accuracy while it lasted but left participants less accurate and more likely to give up once it was removed; the peer-reviewed paper, presented at COLM, argues assistants should scaffold competence rather than just hand over answers.