Review

ChatGPT Improves Polish, Critical Thinking Improves Originality: Lessons from 1,000+ Students

OpenAI published results on August 27, 2026 from a randomized experiment conducted with researchers at Bocconi University. More than 1,000 first-year students completed a real marketing case and were assigned to one of four conditions: access to ChatGPT using GPT‑4o, causal-reasoning training, both, or neither. ChatGPT access raised human-graded work by almost a full point on a five-point rubric and produced more ideas, clearer logic, and work that looked more like expert recommendations. Causal-reasoning training did not raise the conventional rubric score, but students produced a wider and more distinctive range of ideas and explained more clearly why proposals might work or fail. The combined condition preserved both types of gains. The study suggests that AI-era education should evaluate more than polished final answers.

# ChatGPT Improves Polish, Critical Thinking Improves Originality: Lessons from 1,000+ Students ## Article Summary OpenAI published results on August 27, 2026 from a randomized experiment conducted with researchers at Bocconi University. More than 1,000 first-year students completed a real marketing case and were assigned to one of four conditions: access to ChatGPT using GPT‑4o, causal-reasoning training, both, or neither. ChatGPT access raised human-graded work by almost a full point on a five-point rubric and produced more ideas, clearer logic, and work that looked more like expert recommendations. Causal-reasoning training did not raise the conventional rubric score, but students produced a wider and more distinctive range of ideas and explained more clearly why proposals might work or fail. The combined condition preserved both types of gains. The study suggests that AI-era education should evaluate more than polished final answers. --- The education debate is often reduced to one question: > Does ChatGPT make students smarter or lazier? That framing is too broad. Learning includes multiple dimensions: - output quality; - originality; - reasoning; - transfer; - reflection; - alternative generation; - error correction. The Bocconi experiment helps separate some of those effects. ## The assignment was a real-world business case More than 1,000 first-year students developed marketing recommendations for the university merchandise store. This is closer to real knowledge work than a multiple-choice test. It resembles tasks found in consulting, marketing, product strategy, and business analysis. ## Four randomized conditions Students were assigned by class period to: ### ChatGPT access Using GPT‑4o. ### Causal-reasoning training A critical-thinking exercise focused on cause and effect and on explaining why a solution might or might not work. ### Both ChatGPT plus the reasoning exercise. ### Control Neither intervention. This structure allows the researchers to separate the effects of AI access and thinking-skills training. ## ChatGPT improved professional quality Students with ChatGPT access scored almost a full point higher on the five-point human-grading rubric. Their submissions included: - more ideas; - clearer logic; - stronger structure; - greater similarity to expert recommendations. A practical interpretation is that AI helped novices close part of the expertise gap in producing a professional-looking response. ## That does not mean the AI completed the entire task Students still had to decide what to ask, evaluate outputs, choose recommendations, and prepare the final submission. The experimental setup is closer to: ```text student → AI as cognitive tool → final work ``` than simple copy-paste automation. ## The critical-thinking result is more surprising Students who completed causal-reasoning training did not score significantly higher on the conventional rubric. If the analysis stopped there, the intervention might look ineffective. But text analysis revealed other benefits. Those students: - generated a broader range of ideas; - produced more distinctive ideas relative to peers; - explained causal mechanisms more clearly; - discussed conditions under which ideas might fail. Their thinking changed in ways the standard rubric did not fully reward. ## Traditional grading may reward exactly what AI is good at A conventional rubric often values: ```text clarity completeness logical organization standard goals ``` Generative AI is increasingly strong at those qualities. But the rubric may underweight: - unusual ideas; - alternative hypotheses; - challenge to assumptions; - counterexamples; - failure conditions. That creates an assessment problem. ## Final-answer-only assignments lose signal Historically, a polished essay could provide evidence about a student’s writing and reasoning ability. When AI can help almost everyone produce polished prose, the final artifact contains less information about the underlying process. The same issue exists in hiring. If every candidate can generate a professional strategy memo, the PDF alone becomes a weak measure of judgment. ## What should be evaluated instead? ### Originality Did the student explore a distinct path? ### Causal reasoning Why should the proposal work? ### Failure conditions When would it fail? ### Process evidence How did the student move from problem to conclusion? ### Alternative comparison Were multiple plausible approaches considered? ### Revision quality Can the student update after receiving contrary evidence? ## Why AI and critical thinking are complementary ChatGPT improved: ```text answer quality organization idea count logical coherence ``` The critical-thinking exercise improved: ```text idea diversity causal explanation assumption testing failure analysis ``` These are overlapping but not identical skills. Students receiving both interventions showed benefits across the broadest set of measures. ## What schools can change The policy choice does not have to be: ```text ban AI ``` versus: ```text allow unrestricted AI ``` A stronger approach is: ```text allow AI + redesign assessment ``` For example, ask students to submit the final proposal and then explain: - which AI suggestions they rejected; - what evidence could change their conclusion; - what assumptions are most fragile; - why one alternative was not selected. An oral defense can probe whether the student actually understands the work. ## Enterprise training faces the same issue Many corporate AI programs focus only on prompting. A stronger capability model is: ```text AI tool skill + domain knowledge + critical thinking + verification ``` AI can help produce structure and alternatives. Employees still need to judge assumptions, validate evidence, and own the decision. ## Performance systems may reward the wrong thing If employees are measured primarily on how polished their reports look, AI can rapidly improve everyone’s apparent performance. But organizations actually care about: - decision quality; - business outcomes; - risk detection; - innovation; - adaptability. Assessment systems need to shift accordingly. ## Consider evaluating the AI-use process A submission could include: ```text initial hypothesis key prompts AI suggestions rejected suggestion reason for rejection final decision ``` The purpose is not surveillance. It is to measure whether meaningful judgment occurred. ## Expert-like does not equal original This is one of the study’s most useful distinctions. ChatGPT-assisted work was more similar to expert responses. That is generally good for professional quality. But convergence toward expert-like answers is not the same as originality. Critical-thinking training improved idea diversity on a different axis. ## A better AI learning workflow ```text think independently first → write initial hypothesis → use AI to expand alternatives → search for counterexamples → compare options → explain causality → produce final answer ``` This is far more educational than: ```text paste assignment into ChatGPT → submit ``` ## AI literacy is broader than prompting A complete capability includes: - prompting; - evaluation; - verification; - counterargument; - causal reasoning; - ownership. The most valuable skill may be learning when the AI’s polished answer is not enough. ## Conclusion The experiment does not support a simple slogan. It does not show that AI automatically eliminates thinking. It also does not show that AI replaces critical-thinking education. A better conclusion is: > AI access and reasoning training improve different dimensions of performance. ChatGPT helped students produce more professional, coherent outputs. Critical-thinking training increased breadth, distinctiveness, and causal reasoning. The bigger educational question is therefore changing from: > Should students use AI? To: > When polished answers become cheap, which human capabilities should assessment reward? That question applies equally to education, hiring, professional training, and knowledge work. For more analysis of AI education, ChatGPT, and the future of knowledge work, visit **Zyentor Picks**: https://www.zyentorpicks.com/.

Disclaimer: Tool features and pricing may change. Please verify with official sources. Some links may contain affiliate codes.