The Grades Went Up and the Knowledge Went Down. A Study of 3.2 Million Problems Caught Both at Once
A ten-year panel of maths learning data found students spending 27 per cent less time on the problems AI can do, scoring better on unsupervised work, and 25 per cent worse when someone was watching.
Outspoken Digest Technology Desk
Tuesday, August 18, 2026/5 min read

Every argument about artificial intelligence in schools has been conducted in the dark. Teachers report that something has changed. Students report that it has not. Surveys ask people to describe their own honesty and receive the answers you would expect.
A paper published this year finally put a number on it, and the number is worse than the anecdotes, because it does not depend on anybody telling the truth about themselves.
What did the study actually measure?
The paper, Faster Completion, Less Learning, by Sina Rismanchian, Hasan Uzun, Jeffrey Matayoshi, Eric Cosyn and Eyad Kurd-Misto, works from a ten-year panel of 3.2 million learning interactions on the ALEKS maths platform, alongside placement assessment data.
The design is the clever part and it is worth understanding, because it is what makes the result hard to argue with.
The researchers split problems into two kinds: those a chatbot can readily do, meaning text-based word problems, and those it cannot, meaning interactive graph-based problems. Both kinds sit on the same platform, are taken by the same students, in the same sessions. If something changed in students generally, or in the platform, or in the curriculum, both kinds should move together.
They did not move together. After ChatGPT's release, time spent on the AI-susceptible problems fell and time on the AI-resistant ones did not. That divergence is the measurement, and no self-report is involved anywhere in it.
How much less time are students spending?
Among college students, learning time on the problems AI can do fell 2.8 per cent per quarter, compounding to a 26.9 per cent decline over eleven quarters.
The age pattern is the part educators should read twice.
- High school students: down 31.3 per cent. The steepest fall of any group.
- College students: down 26.9 per cent.
- Middle school students: down 9.0 per cent.
- Grade 5 students: no detectable change at all.
The effect scales with age, which is to say it scales with independence and with unsupervised homework. Ten year olds mostly do their work where an adult can see them. Sixteen year olds do not.
The finding that should worry everyone
Here is where the paper stops being interesting and starts being alarming.
The researchers ran the same statistical model on two kinds of assessment: proctored questions, randomly assigned, with someone watching, and ordinary unsupervised assessment.
On the proctored items, the odds of a correct answer declined by 25 per cent cumulatively. On the unproctored assessments, performance went up substantially over the same period.
Read that again, because a great deal follows from it. The same students, over the same years, got visibly better at the work that is graded without supervision and measurably worse at the work that is supervised. Grades rose. Knowledge fell. And every institution in the chain was reading the grades.
The authors call this a population-level indicator of cognitive surrender. It is also, more prosaically, a measurement system that has stopped measuring and started flattering.
One further detail closes the loop: among college students, the entire post-ChatGPT divergence vanishes under proctoring. Whatever is happening is not a general efficiency gain, because an efficiency gain would show up when someone is watching too. It only appears when nobody is.
Can AI detectors fix this?
No, and schools relying on them are exposing themselves to something worse than the original problem.
A study of fourteen AI detection tools found false positive rates as high as 50 per cent and false negative rates as high as 100 per cent. Roughly 20 per cent of AI-generated text was classified as human-written. That rose to around 52 per cent when the text had been lightly edited by hand, and about 71 per cent when it had been machine-paraphrased.
Consider what those two error rates mean in a real school. The false negatives mean any student willing to spend two minutes rewording gets through. The false positives mean honest students get accused, with no way to prove a negative, on the word of a tool whose own error rate would be unacceptable in any other part of the institution.
Detection is not merely ineffective. It transfers the cost of the school's measurement problem onto the students least equipped to fight back, and it does so with a machine that is wrong often enough to be a coin toss on paraphrased text.
What is actually being lost?
Not facts. Facts were never the scarce thing, and anyone who claims schools existed to transmit them has not looked at a search engine since 2004.
What is being lost is the struggle. The productive difficulty of sitting with a problem you cannot immediately solve is not an unfortunate side effect of learning. On the evidence of this paper, it substantially is the learning. Remove it and the work still gets finished, faster and to a higher apparent standard, while the thing the work was supposed to build does not get built.
The supporting data points the same way. Students who lean heavily on these tools score around 6.71 points lower on a hundred point scale than those who do not. Weaker students see the largest apparent grade improvements, which sounds like equity and is closer to the opposite: the pupils with the least secure foundations are getting the most convincing illusion of progress.
What actually works?
The paper contains its own answer, and it is not a policy anyone will enjoy implementing.
Proctoring erased the effect completely. Not reduced it. Erased it. That is the closest thing to a controlled experiment this field has produced, and it points somewhere specific: the problem is not the technology, it is the assessment. We have spent twenty years moving evaluation out of the room and into the home, and we did it while assuming the student was alone.
That assumption is now false, and 59 per cent of academic leaders already report that cheating has risen, while 59 per cent of American teenagers say students at their school use chatbots to cheat at least somewhat often. Around one in ten say they use them for all or most of their schoolwork.
The institutional response so far has mostly been to buy detection software and write policies. The evidence says the response should be to decide what must be verified, verify that in a room, and stop pretending an unsupervised essay measures anything about a person.
What that implies for how schools should be organised is a much larger argument, and one we take up in our companion piece on why the old schooling model is now the thing doing the damage. The evidence here is not an argument for banning these tools. It is an argument for being honest about what our current measurements are worth, which is currently very little.
Published in The Outspoken Digest
Editorial desk
Outspoken Digest Technology DeskSoftware, hardware, artificial intelligence and what they change for everyone else.
Newsletter
The Digest, in your inbox
One edition, sent when it is ready. No noise, and your address is never passed on.
Read Next
More Technology →
A Middle East Data Breach Now Costs 8 Million Dollars, and a Quarter of Them Are AI-Enabled
Aug 18, 2026/3 min read

Alibaba's Open Models Passed Three Billion Downloads. Distribution Is the Whole Story
Aug 17, 2026/3 min read


There Are Now Computers Doing Real Work in Orbit, and a Long Way to Go
Aug 15, 2026/4 min read