← All posts / Research

The Skills That Earn Top Grades Are the Ones AI Can Fake Best: Inside Bocconi's 1,053-Student GPT-4o Experiment

A randomized trial at Bocconi University split 1,053 freshmen into four arms — GPT-4o, causal-reasoning training, both, or neither. ChatGPT lifted scores by 0.86 points, but the grading rubric rewarded conventionality and penalized originality — the exact skills that differentiate human thinkers.

The Skills That Earn Top Grades Are the Ones AI Can Fake Best: Inside Bocconi's 1,053-Student GPT-4o Experiment

What does a top grade actually measure? A randomized controlled trial at Bocconi University in Milan, run in collaboration with OpenAI and published in late August 2026, offers the most precise answer yet — and it is an uncomfortable one. The skills that earn the highest marks on standard academic assignments turn out to be precisely the skills a frontier LLM can most convincingly imitate.

The study, formally titled Training novices to think, or giving them LLMs? Evidence from an RCT, was led by Alfonso Gambardella and Myriam Mariani of Bocconi’s Department of Management and Technology, with co-authors from Duke, Berkeley, SDA Bocconi, and the ION Management Science Lab — several of them OpenAI-affiliated researchers. It is one of the largest classroom randomizations of generative AI ever run.

The experiment

In November 2025, 13 sections of an introductory management course were randomly split into four arms: a control group, a group that received a short lesson in causal reasoning, a group with access to ChatGPT Edu running GPT-4o, and a group that got both. All 1,053 first-year students in economics, finance, and management faced the same task: propose ways to increase awareness and use of Bocconi merchandise among alumni, in up to 180 words. It is a classic marketing-of-new-products assignment — and, as the authors note, exactly the kind of well-structured problem LLMs are built for.

The causal-reasoning lesson was deliberately compact. It walked students through coherent causal logic, falsifiability, and how a proposed action might actually lead to a desired outcome — the mechanics of “why would this work, and under what conditions would it fail?”

Grading was blind: human graders scored the answers on a 1-to-5 scale without knowing which arm each submission came from. Several secondary measures, like causal-reasoning quality and idea diversity, were evaluated with OpenAI and Anthropic models.

What GPT-4o did to the numbers

Access to ChatGPT was close to a full grade-band jump: scores rose by roughly 0.86 points against a control-group mean of 2.09 — a 41% improvement. The AI-assisted answers contained about two more ideas on average, were more logically coherent, and converged toward the recommendations of subject-matter experts. And it was not just polish: even after controlling for argumentation quality, idea count, idea diversity, and text properties, a measurable residual advantage remained, which the authors attribute to higher content quality — not greater student knowledge.

The causal-reasoning lesson, by contrast, did not raise traditional scores at all. Work from that arm actually scored slightly worse on average. But those students more often explained why a proposed action should work and under what conditions it might fail, and they generated ideas that diverged from what their peers wrote. Combining the lesson with GPT-4o added nothing to traditional scores, though the two approaches complemented each other on causal-reasoning markers, and the idea diversity from the lesson group held up.

The rubric is the problem

The most provocative finding sits in the correlation structure of the grades. More ideas, more coherent arguments, and greater within-answer diversity all correlated with higher scores. But stronger falsifiability, more detailed mechanism explanations, and greater divergence from other students’ ideas correlated with lower scores.

The grading rubric, in other words, rewarded well-structured answers that stayed inside the expected solution space — and quietly penalized the very markers of independent thinking the causal-reasoning lesson was designed to produce. The authors’ conclusion is blunt: diversity and originality need to be explicitly built into grading criteria if they are supposed to count. Current grading systems measure polish, structure, and completeness — the exact dimensions where AI assistance shines — while remaining blind to learning and understanding.

That misalignment is what makes AI such an effective cheating instrument. When the instrument of assessment rewards the things a model can fabricate, the final product no longer says much about what the student actually understands.

What the study cannot show

The authors are careful about limits. There was no follow-up test where students had to demonstrate retention without ChatGPT, so whether the GPT advantage reflected genuine learning or merely better output remains open. The experiment covered freshmen at a single university on a narrow marketing task, and randomization happened across 13 class sections rather than individual students. There is also an obvious conflict-of-interest elephant in the room: OpenAI provided the technology being studied, several authors work or worked there, and its VP of Education, Leah Belsky, supplied the accompanying quote about “pairing new technology with thoughtful teaching.”

The broader literature is less sanguine. A study of more than 500,000 US college grades found top grades rose most after ChatGPT’s launch in writing- and programming-heavy courses with large homework components. Controlled experiments showed participants performed worse than controls after AI was taken away — steepest among those who had used it to fetch direct answers. And a 30-month study of over 26,000 Chinese students found homework improved while exam scores dropped, with long-term entrance-exam results 18–24% lower — except among students who spent as much time on homework as non-users despite having access.

Why it matters beyond the classroom

The implications reach into every organization that assesses people against codified criteria. If recruitment, promotions, pitches, and project approvals are scored on predictable rubrics, AI can rapidly close the gap between junior employees and experts — because producing the expected answer is now the cheap part. Firms and universities that actually want independent thinking will need assessment systems that structurally reward it: originality, falsifiable reasoning, and consideration of multiple approaches, not just the most polished conventional answer.

The Bocconi team frames the result not as a verdict against AI but as a design challenge: keep teaching basic principles because LLMs exist, since causal-reasoning training and AI use complement rather than cancel each other. The uncomfortable question the study leaves behind is aimed squarely at the assessors: in a world where machines can produce the expected answer on demand, are you sure your rubric is still measuring the thing you think it’s measuring?