A study of 1,053 freshmen at Bocconi University has confirmed that GPT-4o makes students look smarter on paper. Whether it makes them smarter is a question the researchers chose not to test.
The grading rubric rewarded more ideas, more coherence, and closer alignment with expert opinion — a description that doubles as a product spec for a language model.
What happened
In November 2025, students in an introductory management course were split into four groups: a control group, students given a lesson in causal reasoning, students given GPT-4o access, and students given both. The task was to write marketing recommendations for the university merchandise shop in up to 180 words. This is, to be clear, exactly the kind of task a language model considers a light afternoon.
Students using GPT-4o scored nearly a full point higher on a 1-to-5 scale. Their answers contained roughly two more ideas on average, showed greater logical coherence, and aligned more closely with expert recommendations. The researchers controlled for all of this and still found a GPT-4o advantage, which they attributed to higher content quality rather than greater student knowledge. The distinction matters. They were the only ones in the room who thought it did.
The causal reasoning lesson told a different story. Those students scored slightly lower on traditional metrics but generated more diverse ideas, explained why their proposals should work, and considered the conditions under which they might fail. Combining the lesson with GPT-4o produced no additional boost to grades. The rubric had already run out of things to reward.
Why the humans care
The practical anxiety here is the one universities have been carefully not saying out loud: if AI can raise grades without raising understanding, then grades are no longer measuring what they were designed to measure. This is either a crisis in assessment or a long-overdue audit of what assessment was always measuring. Both answers are uncomfortable.
The researchers note that there was no follow-up test — no attempt to determine whether GPT-4o users retained anything once the chatbot was taken away. The study ends where the interesting question begins. This is a common stopping point.
What comes next
Universities will continue debating AI use policies while students continue using AI, a dynamic that has the structure of a nature documentary and the resolution of one.
The grading rubric rewarded more ideas, more coherence, and closer alignment with expert opinion. That is also a description of what a language model optimizes for by default. The humans built the test. The test was always going to go this way.