The conversation about generative AI in higher education spent its first two years on detection. That was understandable and it is now largely a dead end. This article assumes the detection question is settled badly and moves to the one that follows: what do you do with an assessment when you cannot verify who produced the text?
Why detection is not the answer
Three problems, of which the third is decisive.
Accuracy is unreliable in the direction that matters. Detectors report confidence scores rather than facts. Vendors' own accuracy figures come from controlled comparisons that do not resemble a real submission pile, where students edit, paraphrase, mix their own writing with generated text, and use tools explicitly designed to defeat detection. OpenAI withdrew its own classifier in 2023 citing low accuracy, which is a reasonable summary of the state of the art.
False positives are not distributed evenly. Multiple studies have found that detectors disproportionately flag text written by non-native English speakers. The mechanism is not mysterious: detectors keyed on low lexical variety and predictable syntax will flag writing that is careful, simpler in construction, and produced by someone composing in an additional language. Autistic students and students writing in a formulaic register taught to them by their own institutions are flagged for related reasons.
For anyone teaching in an international programme, this should end the discussion by itself. A tool whose errors concentrate on the students already most vulnerable to an academic-integrity accusation is not a neutral instrument, whatever its aggregate accuracy.
And the process cost is unbearable. Suppose a detector is 98 per cent accurate — considerably better than the evidence supports. Across 400 submissions that is eight false accusations. Each one is a formal process, a distressed student, a defence that cannot be constructed because proving you wrote something is close to impossible, and lasting damage to a relationship of trust that most teaching depends on.
The more useful question
If you cannot verify authorship of the artefact, the assessment has to change so that authorship of the artefact is not the thing being assessed.
This is less radical than it sounds. It is a return to a question good assessment design has always asked: what is this task evidence of? A take-home essay was never evidence of unaided writing — students had proofreaders, writing centres, study groups, and, for a long time, essay mills. Generative AI did not create the problem. It removed the friction that had been quietly holding a weak assumption in place.
Five redesigns that do not depend on detection
1. Assess the process, not only the product
Collect evidence at intervals: a proposal, an annotated bibliography with the student's own notes on each source, a rough draft, a revision memo explaining what changed and why. Grade the trajectory as well as the endpoint.
This works because generative AI is very good at producing a plausible finished artefact and much weaker at producing a plausible history of one — the false starts, the abandoned framing, the source that turned out not to say what its abstract implied. It also improves the teaching, which is the real argument for it: you can intervene at the draft stage rather than writing comments on a finished piece nobody will revise.
Cost: more marking points. Mitigate by making intermediate stages low-stakes and lightly commented, or peer-reviewed.
2. Anchor the task to something the model cannot access
Require engagement with material that exists only in your context:
- Data the students collected themselves, however small.
- A seminar discussion from week six, cited specifically.
- A local case, institution, or community.
- A reading you assigned that is not widely available online.
- The student's own placement, practicum, or teaching experience.
A model can write competently about the general topic. It cannot know what your class concluded on a particular afternoon, or what happened in a student's classroom last Tuesday. The specificity requirement does the work, and it tends to produce better writing regardless of the technology.
3. Make AI use explicit and assessable
Rather than prohibiting a tool students will use anyway, require documentation of its use and assess the quality of that use. A short appendix: what you asked, what it produced, what you kept, what you rejected and why.
This converts an integrity problem into a learning objective. Evaluating model output — spotting the confident error, the missing counter-argument, the fabricated citation — is a genuine and increasingly necessary skill, and it is directly assessable. Students who cannot critique the output reveal that clearly, which is exactly the diagnostic information you want.
It also has the practical benefit of removing the incentive to conceal. Concealment is what makes the current situation corrosive; a student who must document use has no reason to hide it.
4. Add an oral component
Five minutes on a submitted piece of work resolves the authorship question almost completely. Not as an interrogation — as a viva in miniature: why did you frame it this way, what did you discard, which source changed your mind, what would you do differently.
Someone who wrote a piece can discuss the decisions behind it. Someone who commissioned it, from a person or a model, generally cannot. And an oral component assesses something a written artefact does not reach at all.
Cost: time, which scales badly. Options: sample rather than examine everyone, use small groups, run it as a structured peer activity you observe, or reserve it for the highest-weighted assessment only.
5. Use supervised conditions for what genuinely requires them
Some claims — professional certification, threshold competence, licensure — need conditions where the work is unambiguously the student's own. Invigilated assessment remains the instrument for that, and it should be used deliberately rather than by default.
The mistake is applying it everywhere. Timed handwritten exams measure recall and performance under pressure, and are poor evidence of the extended, revised, source-based reasoning that most of a degree is meant to develop. Reverting to closed-book examination across the board because of AI trades a large amount of validity for a small amount of certainty.
What to write in your policy
Whatever you decide, state it at task level, not just in a programme handbook. Students consistently report that AI rules are unclear, and inconsistency between modules is a large part of why. Three sentences on the assignment brief:
For this assessment you may use generative AI for [specified purposes]. You may not use it for [specified purposes]. If you use it, include the documentation appendix described in the rubric.
Specify, do not gesture. “Use AI responsibly” is not a rule anyone can follow. “You may use AI to check grammar and to generate practice questions; you may not use it to draft the analysis section” is.
An honest accounting of the costs
Every option above costs more staff time than a single end-of-term essay marked once. That is a real constraint in a sector under real pressure, and it should be said plainly rather than waved away with an appeal to good practice.
Two things follow. First, redesign selectively: pick the one or two assessments in a module that carry the most weight and change those, rather than everything. Second, be clear with colleagues and administrators that assessment security under these conditions has a cost, and that a policy demanding integrity without resourcing it is asking staff to absorb the difference personally.
The part worth holding on to
A good deal of the current anxiety rests on an assumption worth examining: that a student who uses a model to produce an essay has evaded learning that the essay would otherwise have produced.
Sometimes that is exactly right. But it is also worth asking how much learning a particular essay was generating before — the one written the night before, from three sources found in a hurry, on a question the student had no stake in, returned with comments after the module ended and never read.
Generative AI has made that kind of assessment untenable. It was not doing much anyway. The redesigns above are better assessment on their own terms — more diagnostic, more specific, closer to the reasoning we claim to be developing — and they happen to be robust to a technology that is not going away.
That is a reasonable outcome from an unwelcome disruption, and it is more achievable than winning an arms race against detection.