Recent meta-analyses and research have shown that reasoning-focused models (OpenAI’s o1 and GPT-4) outperform medical students on examinations, clearly and systematically beating the median score and, in most cases, the highest-performing humans’ scores too. The assessments in question were multiple-choice tests, the ultimate symbol of mass assessment.
Last year, the Stanford Human-Centered AI Index tracking data showed something else: when AI is given more time to “think” on examination-style tasks, previous human benchmarks are destroyed. On an International Mathematical Olympiad qualifying exam, an optimized AI model with more “reasoning” outperformed standard AI models (which already outperform humans) by over 60%. In other words, AI will outstrip humans even more on classical assessments if given more time.
Should it worry us that AI can outstrip humans easily on examinations and that enhanced AI will outstrip us even more?
To me, the answer is yes, but not so much because of AI and more because of the type of assessment we have been using to sort students at exit and entry matriculation points for centuries.
In my book Changing Assessment, I explain the main problems with high-stakes, narrowing assessments. Essentially, authentic learning and genuine critical or creative thinking are superseded by timed performance and the ability to cope under a certain, artificial type of pressure. Furthermore, sociologically advantaged students almost always do better, suggesting that examinations are not actually that meritocratic (wealthier students use extra tutors and have access to more test-taking-engineered environments). Finally, the massive scale and standardized nature of these assessments relegate most of the learning outcomes to the lower levels of Bloom’s Taxonomy, favoring knowledge retention and proof of understanding over judgment and analysis. Hardly any analyses of 21st-century competencies see these as particularly valid for the needs of the future.
Now that we know that AI can outperform humans on examinations, what exactly are we examining for? The deeper mechanism behind examinations has always been to classify individuals and groups, which fits the financial model of late capitalism built on scarcity rather than general distribution: the idea is for some to come out on top and others at the bottom. But AI now comes out on top, so where does that lead us in an already somewhat broken model?
The answer is, in my mind, actually much simpler than we might think, and it applies to all assessments, not just examinations.
First, much more emphasis is needed on oracy, discussion, and interview-type assessments. AI cannot be used in these, and they are far more indicative of the type of life-worthy assessment we should be developing in young people, preparing them for job interviews, presentations, and social exchanges.
Second, if we are to test written production, knowing that students can and will use AI (perhaps even inadvertently given its ubiquity), assessments need to take this into account and allow students to use AI openly, to cite that usage openly, and to integrate their own thinking into the work. I asked AI to give examples of what an AI-integrated examination question might look like, insisting that such an assessment push students to think for themselves and not just to rely on AI (which is a slightly ironic scenario), and this is what it came up with:
| Discipline | Traditional Question | AI-Integrated Redesign |
| History & Political Science | Analyze the causes of the French Revolution, focusing on the roles of the three estates. | Generate a royalist speech using AI, critique its historical blind spots/biases, and rewrite sections to reflect a more accurate, deeply nuanced historical perspective. |
| Literature & Creative Writing | Compare and contrast the themes of isolation in Mary Shelley’s Frankenstein and Kafka’s The Metamorphosis. | Have AI build a movie plot based on these themes, analyze the generic tropes it generates, and write a unique character arc that intentionally subverts those clichés. |
| Environmental Science & Economics | What are the economic benefits and environmental drawbacks of transitioning a city to 100% solar energy? | Generate a standard solar transition master plan using AI, then write an adaptive amendment to salvage the plan after introducing a chaotic, real-world “Black Swan” crisis. |
What’s interesting is that each task is much higher up the rungs of critical and creative thinking than traditional assessments, even though (and this is my human critique of AI’s generated solutions), with access to AI, the AI-integrated responses could potentially be done uniquely using AI. On this, at some point, there must be faith in the test-taker’s deontological code, and this should be made clear in the school’s academic honesty policy. Presuming that this is the case, these answers will show greater cognitive application than traditional examination items and also allow students to use AI in formulating their responses, which is something that is happening in the world of work already.
If AI can pass the examination, it’s not AI that should be the focus of our attention, but the way we design the exam. This is a focal point for all educators designing assessments in today’s fast-moving era of VUCA for the long-overdue valorization of higher-order thinking skills.




















