Exam moderators use AI-generated writing to assess literacy standards

The UK Department for Education is testing the use of AI to generate writing samples used in assessing literacy standards, saying it reduces annual costs by 95 per cent.

Por El Medio Oriente
16 de agosto de 2026
Several primary school children seated at desks writing in notebooks during a lesson in a school classroom.
Primary school students work on their writing assignments in a classroom. The UK is experimenting with AI-generated writing samples to assess literacy standards in schools. (New Scientist)
3 min de lectura
Tamaño del texto

The UK Department for Education (DfE) is testing the use of AI to produce writing samples used in assessing literacy standards of students as they progress between primary and secondary school. According to the department, this method reduces annual costs by 95 per cent.

Around 2,000 moderators are tasked with checking the standard of students' writing at the end of Key Stage 2 (KS2), a critical point in the educational journey of children in England and Wales. To ensure that the grades assigned by teachers are standardised, moderators verify students' work by comparing it against writing samples typically obtained from real work produced by schoolchildren. This process costs around £100,000 per year.

From this year onwards and into next year, some moderators will be presented with outputs generated by AI created by ChatGPT's GPT-5 model. The department says GPT-5 is currently being used to create three collections of writing for a standardisation exercise, while a full exercise in 2026-27 and 2027-28 will continue to contain written work by children obtained under the previous contract.

Around 20 expert moderation managers from local authorities will review the AI-derived material to verify its authenticity before it is used. The department plans to decide in spring 2027 whether to continue producing samples in this way or to revert to acquiring them from an external provider. Its transparency register does not list a formal impact assessment for the system.

However, the use of AI in a key part of educational assessment has raised questions. Rebecca Clarkson, a researcher at Anglia Ruskin University in the UK, said that having a writing sample that is not from a real child "creates a philosophical and ethical problem". "The concern is that AI might change our perspective of what is acceptable in writing, or what is a good standard, or something like that," she said.

Clarkson, who studies assessment and moderation of writing in KS2, has heard from several moderators who have expressed concerns about the possible use of AI-generated samples. She said their criticisms centre "essentially on the fact that it is not writing from real children". She added that the consequences of this are impossible to predict: "Is it a concern if we don't use writing from real children? Who knows. I suppose we won't know what the impact is until we can measure it".

The use of AI in this way should not affect which secondary schools students attend, since KS2 writing assessments take place after students have been assigned to their new schools. However, how well children match up against the synthetic samples could determine the level of support they receive when entering their next school.

Since AI tends to produce flat or generic writing, there is concern that children who speak English as a second language or who are neurodivergent and therefore may not use the same language as the majority will be overlooked. The DfE's own assessment of potential risks notes that "language models produce texts that frequently exclude atypical vocabulary and sentence structures that might be used by neurodivergent students or non-native English speakers", something it plans to mitigate through "thorough review".

The assessment also notes that AI-generated outputs will be "extensively reviewed and edited" and will be subject to review before being used in standardisation processes. Jo-Anne Baird, a researcher at the University of Oxford, said: "The use of AI is seeping into all aspects of our lives and without proper governance, we do not know what the implications will be. In assessment, there is concern that there will be a feedback effect, with AI-generated materials becoming the standards that teachers try to get students to emulate. This could lead us to strange places".

Exam moderators use AI-generated writing to assess literacy | El Medio Oriente