Affiliated researcher Sam Gilbert argues that policymakers should take the potential link between declining educational performance and the use of AI chatbots seriously.

Every three years, the OECD’s Programme for International Student Assessment (PISA) tests 15-year-olds around the world. The latest results make for grim reading.
A sharp fall in PISA test scores in 2022 was attributed to the disruptive impact of the Covid pandemic on schooling. But in 2025 performance declined further, with scores in science, mathematics and reading, down on 2022 by 5, 11, and 16 points respectively.

Notably these are the first results to be published since generative AI tools became widely available. ChatGPT was launched in November 2022, and grew to around 700 million weekly active users by the time students were sitting the 2025 tests. Schoolwork is one of ChatGPT’s main use-cases, accounting for around 18% of all messages sent during the 2025 PISA test window.

PISA surveyed students about their use of AI and found that roughly two-thirds use AI chatbots for schoolwork-related tasks, with more than 10% using them “every day or almost every day”. Students who said they use AI chatbots “to summarise a text I had read”, “to conduct preliminary research on a new topic”, or “to draft texts for writing assignments” performed worse in science than students who said they never use AI tools. The gap in performance is material: around 20 points, equivalent to a year’s worth of schooling, even after adjusting for socioeconomic status. Only students who said they used AI chatbots “about once or twice a week” “to help me learn” bucked the trend.

The pattern holds at country level, when the rank of ChatGPT use per capita (according to OpenAI) is plotted against the rank of decline in PISA test scores. In general, the countries in which ChatGPT is used the most heavily saw the biggest declines in student performance between 2022 and 2025. The correlation between ChatGPT use and PISA test score decline is moderate in mathematics and reading (Spearman ρ = 0.406 and 0.358 respectively), and weak-to-moderate in science (Spearman ρ = 0.281).
One objection might be that ChatGPT penetration is simply a proxy for national income, and that richer countries have less room for improvement in test scores than poorer ones. However, it is not that students in the Netherlands, Denmark, or Singapore failed to improve – it is that their performance went backwards in absolute terms.

In itself, this analysis is not proof that poor PISA scores in 2025 were caused by students’ adoption of ChatGPT, but it does align with academic studies examining the relationship between AI use and student performance.
Fan et al 2024 found that university students who were given access to ChatGPT produced better essays than their peers, but showed no improvement in knowledge when tested later, and performed worse on some measures, leading the authors to posit that relying on ChatGPT leads to “metacognitive laziness”. Oakley et al 2026 offered a neuroscience-based explanation for recent declines in IQ, attributing it to widespread “cognitive offloading” onto AI and other digital tools.
There are some indications that AI tools need not be detrimental to learning. A field experiment involving around 1,000 secondary school students in Turkey found that students with access to a ChatGPT-based maths tutor performed 17% worse than their peers when the access was removed – but that a version of the AI tutor with built-in learning guardrails largely mitigated these effects (Bastani et al 2025).
However, by default, ChatGPT, Gemini, Claude and other AI chatbots have no such guardrails. If a student asks them to solve a quadratic equation, label a diagram of a cell, or write an essay on Gladstone’s foreign policy, that is what they will do. They are currently accessible to every student with a smartphone, for free, at all times. Even those of us who are optimistic about the potential for AI in education would struggle to argue that this kind of unrestricted access is beneficial.
How should policymakers respond? In England there is a case for doing nothing. Cameron-era educational reforms that emphasised knowledge and exams (as opposed to skills and coursework) have been vindicated by PISA scores that have remained stable since 2012, while many other countries have declined sharply. Education systems in Wales, Scotland, Northern Ireland and beyond might look to this formula, in part because it seems to mitigate the downsides of unrestricted AI access.
Another possibility might be to broaden the scope of forthcoming social media legislation. The UK government intends to ban most social media platforms for under 16-year-olds, and has indicated that the legislation will include new measures promoting safe use of AI chatbots. These measures are targeted at different AI harms, such as children becoming emotionally dependent on AI companions or receiving problematic mental health advice. But policymakers might consider introducing a requirement that providers of AI chatbots route schoolwork-related requests to guardrailed versions of the kind tested by Bastani et al. As well as benefitting students’ learning, this could also encourage the development of a market for purpose-built educational AI chatbots.
At present, the evidence that AI chatbots degrade students’ performance is not conclusive, but waiting for causal proof before acting risks letting down yet more cohorts.
The views and opinions expressed in this post are those of the author(s) and not necessarily those of the Bennett School of Public Policy.