AI Education

OpenAI study finds ChatGPT and critical-thinking training improve student work in different ways

OpenAI reported results from a randomized Bocconi University experiment showing ChatGPT access improved polish and coherence while causal-reasoning training increased idea variety.

Published Updated
OpenAIChatGPTAI Education

OpenAI has published results from a randomized education experiment suggesting that ChatGPT access and critical-thinking training improve student work in different but complementary ways. The company posted the findings on August 27, 2026, describing research conducted with Bocconi University and OpenAI Economic Research involving more than 1,000 first-year undergraduate students.

The experiment asked students to work on a real-world business case: developing marketing recommendations for Bocconi University’s merchandise store. Students were randomly assigned by class period into four groups. One group received access to ChatGPT using GPT-4o, one completed training in causal reasoning, one received both, and one received neither. The causal-reasoning exercise was not an AI tutorial. It taught students to connect cause and effect, examine why an idea might work, and consider when it might fail through a game-like exercise with examples, questions and feedback.

Researchers evaluated the submissions in two ways. Human graders used a five-point rubric focused on how well the recommendations addressed standard marketing goals. Separately, automated text analysis measured the number and variety of ideas, evidence of causal reasoning and similarity to recommendations written by experts. That dual approach is important because a polished answer and an original answer are not always the same thing, especially in classrooms where AI tools can make surface quality easier to achieve.

The ChatGPT effect was clear on the grading rubric. Students with access to ChatGPT scored almost a full point higher on the five-point scale. Their work contained more ideas, clearer logic and stronger similarity to expert recommendations. OpenAI emphasized that students were not simply handing assignments to the model. They still had to decide what to ask, judge the responses and choose what to include. The result is closer to an expertise bridge: AI helped novices produce work that resembled more professional analysis.

The critical-thinking exercise produced a different result. Students who completed it were better at explaining why ideas might work and when they might fail, but the traditional rubric did not give them higher scores. Text analysis showed that these students generated a wider range of ideas that were more distinct from their peers’ submissions. That finding points to a measurement problem in education. If grading rewards coherent conventional answers more than intellectual variety, schools may miss some of the skills they say they want students to build.

The students who received both ChatGPT access and causal-reasoning training showed the broadest set of gains. Their idea variety resembled the group that received only the thinking exercise, while their rubric scores and number of ideas resembled the group with only ChatGPT access. Their work also showed stronger logical coherence and more evidence of questioning assumptions. The study therefore pushes against a simple debate in which schools must choose between AI use and independent thinking.

The practical implication is that AI in education should be treated as a design problem, not only a permissions problem. If students can use AI to produce polished work, final answers alone reveal less about what they understand. Assignments may need to reward originality, reasoning process, comparison of alternatives and the ability to defend choices. ChatGPT can help close an expertise gap, but critical-thinking training still matters because it changes how students frame problems before and after the model responds.