A landmark randomized experiment at Bocconi University found that ChatGPT access and critical-thinking training improve student work in different ways — and that students who got both came out ahead on the widest range of measures.
The debate over AI in the classroom tends to get framed as an either/or: either students use AI and shortcut their development, or they learn to think independently. A new study from researchers at Bocconi University, conducted in collaboration with OpenAI Economic Research, offers the most rigorous evidence yet that this framing is wrong — and that the real opportunity lies in combining both.
Published Aug. 2026, the study ran a randomized controlled experiment with more than 1,000 first-year undergraduates at Bocconi, one of Europe’s leading business schools. Students were assigned to one of four groups: access to ChatGPT (specifically GPT-4o), a causal reasoning training exercise, both, or neither. The assignment was grounded and practical — develop a set of marketing recommendations for the university’s own merchandise store. It’s the kind of open-ended, real-world task that shows up constantly in internships and entry-level jobs.
What ChatGPT Actually Did for Student Work
Students with access to GPT-4o scored nearly a full point higher on a five-point rubric graded by trained human evaluators. Their submissions included more ideas, followed clearer internal logic, and more closely resembled the recommendations produced by three domain experts brought in as benchmarks. In short, the AI helped novices produce work that read like it came from someone with more experience.
Critically, students weren’t just copy-pasting ChatGPT outputs. They still had to decide what to prompt, evaluate what came back, and make editorial judgments about what to include. The cognitive load didn’t disappear — it shifted. That distinction matters both ethically and practically: the students doing the work were still the ones steering it.
What Critical Thinking Training Did — and What the Rubric Missed
The causal reasoning training produced a more surprising result. The exercise — built around a game, worked examples, questions and feedback, and entirely unrelated to AI — did not raise rubric scores at all. By the standard grading metric, it looked like it did nothing.
But automated text analysis told a different story. Students who completed the training produced a significantly wider variety of ideas. Their submissions were more distinct from their peers’ — a signal of genuine originality rather than convergence on the most obvious answers. They were also better at explaining why a strategy might work and, just as importantly, when it might fail.
That’s a gap worth naming explicitly: a conventional rubric rewarded polish and correctness but was blind to originality. The training sharpened something the assessment couldn’t see.
The Students Who Got Both Did Best Across the Board
Students in the combined group — ChatGPT access plus causal reasoning training — showed gains across the widest range of measures. Their rubric scores matched those of the ChatGPT-only group. Their idea variety matched the training-only group. And their work showed stronger logical coherence and more evidence of questioning assumptions than either group alone.
“The takeaway: AI helped students make their answers better. Critical-thinking training helped make their ideas broader. The two play complementary roles in preparing students for the future,” concluded OpenAI.
The study’s randomized design is what gives these conclusions real weight. Most research on AI in education relies on surveys, self-reports, or observational data — all of which struggle to separate correlation from cause. A properly randomized experiment, at the scale of 1,000-plus students, is rare, which is why this one is drawing attention beyond just the AI industry.
Where This Lands in a Crowded Market
The study arrives at a moment when every major AI lab is competing on pedagogy, not just raw capability. OpenAI’s ChatGPT Study Mode is designed to guide students through problems step-by-step, functioning more like a Socratic tutor than an answer machine. It competes directly with Khan Academy’s Khanmigo and Google Gemini’s Guided Learning mode, which takes a more structured, course-like approach with visual explanations and progress tracking. Google has also partnered with Pearson and other educational publishers to embed Gemini-powered tools directly into institutional platforms.
Against that backdrop, the Bocconi study functions as something rare in the education AI wars: controlled experimental evidence. While competitors are largely making claims based on product design philosophy or user satisfaction, this study offers data on what actually happened to student output when GPT-4o was introduced under experimental conditions. For OpenAI, that’s a meaningful differentiator in conversations with universities and education policymakers.
What This Means If You’re a Student Right Now
There are two practical takeaways here that go beyond the research paper itself.
First, using ChatGPT on assignments isn’t inherently undermining your own growth — but only if you stay in the driver’s seat. The students in this study who benefited weren’t passive consumers of AI output. They were making decisions: what to ask, what to accept, what to revise. That active role is the difference between building skill and outsourcing it.
Second, and more urgently: be aware of the grading blind spot this study exposed. As AI makes polished, well-structured answers easier to produce, traditional rubrics become less reliable signals of who actually thinks well. The originality and independent reasoning that a conventional grade won’t fully credit are often exactly what distinguishes strong candidates in competitive hiring and graduate admissions. If your coursework doesn’t reward that kind of thinking, it’s worth cultivating it anyway — through debate, case competitions, causal reasoning practice, or any structured exercise that pushes you to explain not just what but why.
The students who combined both tools in this study didn’t just do better on the assignment. They came out with a broader toolkit. In a job market where AI is compressing the value of generic, polished output, that combination is exactly what gives you a durable edge.
Source: OpenAI
Additional research sources
- https://blockchain.news/news/chatgpt-critical-thinking-student-performance
- https://www.medialaws.eu/training-the-mind-in-the-age-of-machines-chatgpt-at-bocconi/
- https://www.sciencedirect.com/science/article/pii/S2666920X26000330
- https://theairankings.com/best-ai-for-students/
- https://www.linewize.com/blog/chatgpt-launches-study-mode-what-k-12-leaders-need-to-know
