Over the past year, a good deal of energy across the sector, mine included, has gone into rethinking assessment for a world where generative artificial intelligence is simply part of how students work.
I have spent much of that time developing a , embedding academic integrity into the design of a task rather than policing it after submission, and this kind of work is spreading fast across the UK. But it only addresses half of the problem.
Assessment design focuses on what we ask students to do. But what about how we mark it? We are through the from an era when marking capacity was the binding constraint, when a module was a sealed unit, a final artefact was the only evidence available, and staff judgement was the only trustworthy source of a grade. AI has transformed all that – but almost nobody is asking what marking should look like now that it has.
One idea I have is to switch from marking modules to marking years. Modular marking treats learning as a series of disconnected certifications rather than what it actually is: a continuous development that happens to pass through several deadlines. AI makes the seams between modules embarrassingly visible because the same student can produce very different quality work depending on how well a particular module brief happened to reward the final artefact – which can be outsourced to AI – over a demonstration of genuine thinking and development.
ߣߣÊÓƵ
A year-long mark, or a programme-level grade, moves us towards something closer to a “value-added” model, which doesn’t look at work in isolation but at how far the student has come over time, through regular checkpoints. This is much harder for AI to replicate because it gives us a holistic picture that’s sustained and deliberate, giving us a bank of evidence to assess at once. If there’s suddenly a change in voice, sophistication or argument between checkpoints with no visible development accompanying it, that would raise a red flag.
There is also a regulatory obstacle, however. Most institutions have regulations and systems designed to make decisions about students’ performance at the level of individual modules, which makes it difficult to introduce a year-level mark or make progression decisions based on a student’s overall performance across the year.
ߣߣÊÓƵ
But even if that obstacle can’t be overcome, we should still consider marking a student’s distance travelled, rather than basing judgement solely on the absolute quality of a final piece against a generic rubric. Currently, two students can submit near-identical polished drafts having done entirely different amounts of thinking to get there. A trajectory model would track the drafts themselves, the false starts, the moments where a student abandoned one argument for a better one. This is something that the evidence suggests AI cannot fake its way into producing convincingly.
That isn’t to say that a trajectory model would punish a strong start by a more able student. The point isn’t that every student needs visible improvement from a weak first attempt. The point is to have evidence that the strong first draft was genuinely their own thinking. A student who arrives at a good answer quickly and then spends the rest of the module testing it, refining it or defending it against counter-arguments is still generating a visible trajectory: it’s just a trajectory of depth rather than of correction.
Nor would a trajectory model replace the final judgement of competence. Employers and professional bodies still need to be sure that a person can do a certain thing, and nothing I’m suggesting stops us setting that bar and marking against it. It just adds evidence that the student passed the bar on their own merit.
Further evidence of that could be provided by vivas. These currently live at the edges of the system, reserved for doctorates and borderline classifications because they are expensive and do not scale. But AI is starting to change the calculus here, too.
ߣߣÊÓƵ
A study used voice AI to run and grade oral exams for two undergraduate cohorts, bringing grading cost down to roughly a dollar per exam – and most students felt the format tested genuine understanding. If a short oral defence becomes cheap to administer and record, the question is why the exception should not become the norm since a live defence of understanding is as close to reliable proof of learning as we can get. A proportion of vivas could still be moderated by a human if doubts remained about the reliability of the AI’s judgement.
By at the mechanical end of marking, AI could also allow staff to oversee the use of student peer review in assessment. Peer review is the engine the entire research system runs on. It is also exactly the skill graduates need in a workplace full of AI-generated first drafts that someone has to evaluate. Yet in undergraduate education we treat it as a formative exercise at best: something students do to each other for practice, while the marks that matter come only from staff. If we had time to oversee it, peer assessment could instead account for a formal proportion of a student’s module mark, moderated and structured against criteria.
Every one of these ideas is pedagogically defensible but administratively terrifying. Yet our obligation is not to defend the old architecture out of habit. It is to ask what marking should look like now that the reasons for its traditional form no longer hold.
Emma Ransome is academic lead for teaching and learning at Birmingham City University.
ߣߣÊÓƵ
Register to continue
Why register?
- Registration is free and only takes a moment
- Once registered, you can read 3 articles a month
- Sign up for our newsletter
Subscribe
Or subscribe for unlimited access to:
- Unlimited access to news, views, insights & reviews
- Digital editions
- Digital access to °ձᷡ’s university and college rankings analysis
Already registered or a current subscriber?







