Contents

The Breaking Point for Higher Education: Student Work Is Getting Better. The Learning Is Not.

Generative AI improves essays, code, and grades faster than universities can work out what a piece of student work now proves — and what stays with its author once the laptop closes.

Picture two students handed the same assignment. Both are bright, both are motivated, both want the top grade — and from there their paths split.

The first loves the subject. She reads the sources, spends a long time building her argument, writes a messy draft, and only then opens ChatGPT — not to replace her thinking but to stress-test it. She asks the model to find the weak spots, raise objections, point to the gaps in her logic, and then she goes back and rewrites the thing herself.

The second student loves the result more than the subject. Study is a means to an end, so he feeds the model the prompt, the grading rubric, and the material he’s gathered, gets back a strong structure, checks the sources, tightens the phrasing, and hands in a nearly flawless paper.

Both papers can earn the top grade. And that is where the problem begins.

The university sees two equally convincing products, and almost none of the process that produced them. One student stretched her own thinking; the other, to put it plainly, borrowed his for the week. On paper their competence looks comparable, even though beyond the text the difference may be enormous.

I look at this from several angles at once. As someone who works in marketing and communications, I already see every day how useful AI can be: I use it, I study what it can do, and I have no intention of pretending the technology can be canceled by administrative decree. But a background in journalism trained me to separate an impressive result from proof, and a convenient explanation from a fact.

There’s another reason the question won’t leave me alone. I have two children. My older daughter faces a serious educational choice in about five years. My younger son will enter the system much later — but already in a world where AI is an ordinary part of school. What should they learn, how should they learn it, and what will a university degree even certify by then?

Not long ago questions like these sounded like futurology. Now they walk into the house with a child who opens a laptop and sits down to do homework.

On August 6, 2026, Denmark offered one of the first policy-level answers. The Ministry of Education announced that major examination papers prepared at home will have to be defended out loud. The first group named directly was the roughly 9,000 students in the HF program — a two-year upper-secondary track that qualifies graduates for university — who each year write a large individual final assignment, the SSO.

This is still upper secondary school, not university. But the two share one problem: the finished paper no longer lets you judge, with any confidence, the independent competence of its author.

Denmark did not commit to a single clean philosophy, though. The same package folded in screen monitoring during exams, network filters, and a recommendation to write more work in school, under supervision. The exact shape of the future defense is still to be worked out with individual institutions.

The result is a telling hybrid: one hand reaches to see competence more clearly, the other to control the production of text more tightly. The international turn from hunting the AI to verifying the human has genuinely begun — but it looks neither uniform nor finished.

When the Work Improves but the Student Doesn’t

Universities are used to grading what they can see and compare: an essay, code, a report, a presentation, an exam answer. Generative AI is superb at improving exactly those outputs. The improvement is easy to see today; the learning shows up later and is far harder to measure.

A recent programming experiment captured the gap in a title that almost writes the diagnosis itself: “Less stress, better scores, same learning.” Researchers compared three groups: one used ChatGPT freely, one worked with a constrained AI tutor, and one went without a generative model at all. AI support helped students complete the task better and lowered their frustration — but the gain in knowledge stayed statistically similar.

The code got better. The understanding didn’t keep pace.

Academic writing showed a similar split. In a randomized experiment, students revised essays with ChatGPT, with a human expert, or with checklists. The ChatGPT group produced stronger texts, yet gained no clear edge in knowledge or transfer, and the study’s authors warned of a risk they called “metacognitive laziness.”

The phrase reads almost like a moral rebuke, though the mechanism it describes is more interesting, and more neutral, than that. The student isn’t necessarily idling: he asks questions, weighs options, edits paragraphs, checks facts, stays busy until late at night. The activity doesn’t vanish. What quietly shifts is the authorship of the thinking.

The model offers a structure before the student has built his own. It notices the missing argument, invents the counterexample, fixes the awkward transition. The student recognizes good reasoning when he sees it on the screen — and that recognition is easy to mistake for understanding.

But recognizing a thought and being able to reproduce it are different abilities. Close the laptop, clear the chat history, ask the person to explain the key argument again from scratch, and the difference can surface fast.

We have lived with cognitive offloading for a long time, and it isn’t a bad thing in itself. Writing freed our memory, the calculator took over part of our arithmetic, the search engine let us stop holding every date in our heads. Civilization advances precisely by handing tools some operations and pointing the freed attention at others.

But every such trade changes the human skill that remains. The calculator cut the volume of routine arithmetic, yet it didn’t write the student’s proof; the search engine found sources, but rarely assembled them into a whole argument on the spot. Generative AI crosses a new line, because it can offload not just memory or retrieval but a large share of the reasoning itself.

Sometimes that’s exactly what’s needed. Not all effort is useful: a bad explanation, a week’s wait for feedback, pointless recopying, and bureaucratic busywork make nobody smarter. Universities produced friction like that long before ChatGPT, and AI may well clear it away.

But there’s another kind of effort — the kind that competence is slowly built from. A hypothesis you form yourself, an error you find, the struggle to recall a principle, the defense of an uncomfortable argument, a problem solved without a ready template: these take time and are sometimes annoying. That is exactly why they work.

AI can spare a student that effort too. Not because it “rots the brain” — there’s no data for a verdict like that yet — but because it makes independent thought optional in the very minutes when it might have turned into knowledge.

So the distinction that matters isn’t between using AI and not using it. It runs between augmenting the human and automating the human.

One IZA working paper, based on an experiment with 211 students, even found an improvement on an independent test right after working with AI, and a week later. But the durable advantage belonged mainly to those who used the model as a partner for inquiry and study. For students who delegated the production of the work to it, the advantage disappeared once the tool was taken away. This isn’t final proof; it’s an early working result. Yet it shows nicely why “did the student use AI?” is too blunt a question. The sharper one is: how much of the thinking did he hand over?

Four Different Claims You Can’t Test the Same Way

Almost every convenient metric belongs to the present moment: the quality of the work, the speed, the error count, the stress level. A researcher can measure them within a semester and get a tidy table. Long-term learning is harder.

What will the person recall in a month? Can she explain the idea without a prompt, carry it into an unfamiliar situation, catch a confidently phrased error from the model? Answering that takes delayed tests, longer observation, and a willingness to wait — while the technology changes faster than the research cycle.

Here it helps to separate four claims that the AI debate keeps blurring together:

  1. Authorship: who produced the work and made the key decisions?
  2. Understanding: can the student explain the result right now?
  3. Retention: what can she reproduce weeks or months from now?
  4. Transfer: can she apply the principle in a new, altered, or unfamiliar situation?

No single check establishes all of them. Version history, drafts, and a prompt log can support a claim about authorship, but not prove deep understanding. A conversation right after submission can test present understanding, but the student may have rehearsed the explanation well. A delayed task without outside help gives evidence of retention. And only a genuinely new problem lets you test transfer in any meaningful way.

So an oral defense doesn’t prove the AI was absent. An immediate conversation doesn’t prove durable learning either. It creates a protected moment in which understanding can be checked. Retention takes time. Transfer takes a new problem.

We have piled up a fair amount of evidence that AI improves today’s task, and remarkably little on what remains with the student a few weeks later. That isn’t proof of harm. It’s proof of uncertainty — in exactly the part universities formally exist to deliver.

Meanwhile, the scaling has already happened. According to the Digital Education Council (DEC) 2026 survey, 88% of students and 77% of faculty use AI. Its use stopped being an experiment long ago, while the evidence for durable learning is still being gathered in experimental mode.

That gap is what strikes me as the real breaking point: the system has already changed the behavior of millions of people before working out what learning it now produces.

Not All AI Leads Away From Learning

It would be too convenient to end on an alarmed note and declare generative models a machine for intellectual softening. The data won’t allow it. In fact, some experiments show that a well-designed AI can teach very well.

In a randomized Harvard study, students worked with a purpose-built AI tutor in physics. They absorbed more material, and the median gain in knowledge came out more than twice as high as in an in-person active-learning class — in less time.

The result is impressive, and it does give grounds for cautious optimism. But this was not a free conversation with an ordinary ChatGPT. Instructors defined the learning goals in advance, experts prepared the content, and the system led the student along a designed path, managed the cognitive load, gave targeted feedback, and refused to let him keep choosing the easiest route.

The authors name the core distinction plainly: an AI chatbot is usually built to help its user, not to help that user learn — and those are not the same goal.

The study carries real limits: it covered two physics sessions, measured the result immediately afterward, and ran no delayed test of retention. So the example doesn’t justify handing chatbots to every student without thought. It shows something more valuable: what matters isn’t access to AI as such, but the mental work the system lets a person skip.

A good tutor removes confusion without canceling the attempt. It can feel less obliging, because it doesn’t hand over the full answer at once — it asks one more question, requests a prediction, points back to a contradiction, and sometimes lets a little frustration do its work. For a product people are used to judging by speed and convenience, that’s almost a paradox: the best educational AI has to know when to refuse to help.

/uploads/ai-tutor-collaborative-learning.jpg

One Tool, Two Different Goals

Back to the two students from the opening. The first is optimizing for understanding, so she turns the model into an intellectual sparring partner: she asks for counterarguments, hunts for missing evidence, tests whether her idea survives being moved into a new situation. The second is optimizing for a measurable result, so he uses AI as production infrastructure — something to absorb the search, the structure, the draft, and the polish.

Both can be equally smart, and neither needs bad intentions. The technology simply amplifies the goal a person brought to it: for one, deeper understanding; for the other, maximally efficient success.

The second student is easy to condemn, and universities even benefit from that explanation. Call him lazy or dishonest, and nothing about the system itself has to change. But that student may have read the system with great accuracy.

The university rewards the final product — so a rational person improves the final product. Employers still demand a degree, tuition is expensive, and the day has not grown extra hours. Why should an ambitious student volunteer to give up a tool that gets him through the formal check faster?

Working in marketing showed me the same law over and over: people optimize the measured metric fairly consistently, and follow the stated goal far less often. A university can talk about critical thinking all it likes, but if the grade is attached to a take-home essay, the student will optimize the take-home essay. AI now produces that signal cheaply.

The student’s responsibility doesn’t vanish. Blind delegation builds brittle competence, and a flawless text can hide empty confidence. But the rules of the game are designed by institutions. After generative AI, they can no longer reward the external result alone and assume they have measured the human.

First, Universities Tried to Catch the Machine

The first reaction was academic integrity. That made sense: ChatGPT arrived mid-year, policies lagged, instructors received unfamiliar texts, administrations needed rules immediately.

Public memory overstates the scale of blanket bans. A UNESCO survey of more than 450 institutions in 2023 found only two cases of a broad ban. A different figure is far more telling: fewer than 10% of the organizations surveyed had any formal guidance, and among universities only 13% did.

The institutions turned out to be not so much authoritarian as unprepared. And in that confusion, detectors looked like an almost perfect technical escape: if one machine made the text, maybe another could expose it.

The hope met reality quickly. Stanford researchers tested seven detectors on TOEFL essays written by humans who were not native English speakers. On average the systems wrongly flagged 61.22% of those papers as generated, and at least one detector suspected AI in 97% of the essays.

This was not a mere technical glitch but a structural unfairness: predictable language became suspect, and international students were placed at disproportionate risk. At the same time the familiar arms race began. Models improved, paraphrasing tools multiplied, and telling human text from machine text with any confidence grew steadily harder.

Some universities drew the practical conclusion. As early as August 2023, Vanderbilt disabled the Turnitin detector, citing false positives, opacity, and doubt about whether reliable automatic detection is even possible. Other institutions later dropped the tool as well.

None of this means violations disappeared, or that any use of AI became acceptable. What changed was the object of trust. Instead of asking “can a program guess where the text came from?”, schools increasingly have to ask a different question: “what observable act would let a person show the competence we need?”

First universities tried to detect the machine. Then some began to redesign the task. The turn was the right one — but it’s a long way from finished.

/uploads/ai-detector-academic-polygraph.jpg

A License Is Not a Reform

Today universities publish policies, hand out access to models, build protected platforms, and add AI-literacy courses. These are real changes, yet the presence of the technology inside the campus doesn’t mean the learning has been transformed.

An institution’s response is more usefully judged by the state of the system than by the number of licenses:

  1. Announced: principles and guidance exist.
  2. Operational: concrete rules apply to students and faculty.
  3. Piloted: real tasks or exams have been rebuilt in trials.
  4. Scaled: the new model runs at the level of a program, an institution, or a national system.
  5. Verified: understanding, retention, or transfer has been measured independently.

Rolling out Copilot or ChatGPT Enterprise is not a rung on that ladder. It’s an infrastructure decision. A university with an expensive platform but the same old take-home essay may be pedagogically behind an institution that has no model of its own but has already changed how it assesses.

The DEC 2026 data shows progress and shallowness at once. Only 19% of faculty reported a significant redesign of their work, another 53% made minor adjustments, and just 20% called their AI integration proactive.

Students feel the gap especially clearly. Forty-three percent met AI in none of their courses, while only 15% saw it across many. Among those who did encounter integration, just 5% felt it genuinely transformed their learning; another 28% reported a deeper grasp of the material. Meanwhile 76% of students received no institutional training in working with AI, and 57% got no clear rules for using it in graded work.

Universities are getting better equipped with AI, but not necessarily better at teaching with it. A license doesn’t solve the pedagogy, a policy doesn’t prove acquired knowledge, and a prompt workshop can’t, by itself, rebuild what a degree means.

The Future Has Already Entered the Exam Room

The Danish move is appealing in its simplicity: if the big paper was written at home, ask the student to defend it. But the international picture is more interesting, because different systems arrived not at a single format but at a shared architectural idea.

In Norway, since the 2023/24 school year, schools have been trialing the langtidsoppgave — a long-term assignment offered as an alternative to the usual exam. Students work on a written product, use sources and tools, receive guidance, and then take part in an oral subject conversation. It’s one of the closest school-level analogues of the “open work with AI plus independent verification” model: the system sees not only the final document but the process, the explanation, and the ability to answer for the decisions.

But the evaluation of the Norwegian pilot declared no victory. Implementation varied noticeably across subjects, which hurt comparability. The terms for using AI and other tools stayed uneven, creating a risk of unfairness. The researchers recommended more experiments and clearer national frameworks, not an immediate promotion of the pilot to universal standard.

In higher education, the most complete answer came from the University of Sydney. Its framework splits assessment into two loops.

The open loop is for learning and realistic work. Here AI and other available tools aren’t fought off: students are taught to use them in research, analysis, product creation, and professional practice.

The secured loop is used at the points where the university needs controlled evidence of individual competence. That might be an in-person written or practical exam, an interactive conversation, an observed action, work on a placement, or questions following an AI-assisted project.

Sydney’s key decision is not to demand an oral defense after every essay. Open tasks live mostly at the level of individual courses, while secured checks are designed at the level of the whole program and placed at strategic checkpoints. The university tries not to re-prove every student in every discipline, but to assemble a reliable enough map of evidence by the time the degree is awarded.

One loop assesses the student’s work with the tools available. The other samples what he can do on his own, under controlled conditions.

That’s subtler than the familiar “write it with AI, then defend it without.” The defense is only one instrument. A future architect might take apart a new system, a programmer might fix code live, a clinician might work through a patient scenario, a linguist might write and edit a text under observation. The method has to fit the ability the degree promises to an employer and to society.

Queensland University of Technology: Can Independent Verification Scale?

The strongest new practical example came from nursing education. At Australia’s Queensland University of Technology (QUT), a take-home written analysis of a clinical case was replaced with an in-person oral assessment structured around ISBAR — a professional format for handing over patient information.

In 2025, about 700 students went through the new model; 740 were enrolled in the course. Each received one of 40 clinical scenarios, took 30 minutes of supervised preparation, and then delivered a situation assessment and a recommendation out loud.

Compared with the previous cohort, the spread of course results held steady, the final written-exam scores improved, no academic-integrity violations were recorded, and the marking took less time and money. It’s a rare demonstration that secured assessment can work outside a small seminar. The QUT study describes a cohort on the scale of a large university stream.

But even this result can’t be turned into proof that the new system is superior. The researchers compared consecutive yearly cohorts, not randomly assigned students. The group’s makeup, the teaching, and outside conditions could all have shifted. And QUT replaced a written task with an oral one, rather than combining full AI-assisted production with a separate secured check.

The QUT case shows that a check like this can be run even for several hundred students. It doesn’t yet show that it improves long-term learning.

An Oral Defense Tests Understanding — but Doesn’t Guarantee Fairness

The obvious answer sits right on the surface: if AI could have written the take-home paper, sit the student down and talk to him. Universities call the format a viva voce — Latin for “with the living voice” — or just a viva: an oral exam or defense in which an examiner questions the student about the submitted work.

The conversation reveals what the finished text can’t. You can ask the student to explain his reasoning, justify a decision, apply the idea to a slightly altered problem — and it becomes clear fairly quickly whether he understands his own work or merely brought in a convincing-looking result.

But an oral defense has a weakness of its own: too much depends on who’s sitting across the table. One examiner asks harder questions, another prompts without meaning to, a third rewards confident delivery. A conversation without shared rules risks measuring not knowledge but eloquence, composure under stress, and luck with the examiner.

A review of 24 studies shows that this randomness can be reduced with a structured defense: comparable questions for everyone, marking against predefined criteria, trained examiners. In individual studies, that format produced markedly more consistent grades than a free conversation.

The one measure the authors could pool in a meta-analysis was students’ attitude to the format: most found the oral defense acceptable. But there wasn’t enough comparable data to draw a general conclusion about the accuracy and reliability of such a check.

So an oral defense can become a useful moment of independent verification. What it doesn’t become is a truth detector: the result still rests on the quality of the questions, the consistency of the criteria, and the training of the people asking.

The Cost of Human Contact

At Newcastle University, 60 final-year students, after a project of roughly 6,000 words, sat a 20-minute defense in front of two instructors. With feedback and moderation, each student took 40 minutes, and the whole process ran to about 80 faculty hours. The author of the case doesn’t call it a defense against AI: it’s more of a deterrent that makes handing the whole mental job to a model less worthwhile. To answer an unexpected question, the student still has to understand his own project.

The diagnostic value of an oral defense comes from adaptivity, follow-up questions, and human judgment. But those are exactly what make the format expensive. Scale is often reached by stripping them out: students record short video answers, get identical questions, and the instructor can no longer probe a weak spot on the spot. The more efficient the pipeline, the less it resembles an actual conversation.

The Subjectivity Doesn’t Disappear

Even a trained examiner remains part of the measuring instrument. In a Monash study, grades depended far more on the individual examiner than on which task version a student got. This is still only a conference abstract about formative team assessment, so the result can’t be carried over to every oral exam.

AI weakened trust in the written paper. Oral assessment doesn’t restore it automatically. A poorly designed defense can simply swap one source of uncertainty for another — and measure confidence, speaking speed, or luck with the examiner instead of knowledge.

There’s an ethical line here too:

A failure to defend the work means the student may need further review. It is not, on its own, proof of misconduct.

A weak answer can come from anxiety, language, communication differences, a badly framed question, or examiner bias. Otherwise universities would only be swapping an unreliable algorithmic detector for unreliable human intuition.

Does Defending the Work Change the Learning Itself?

There’s an appealing hypothesis: if a student knows he’ll later have to answer without AI, he’ll work differently on the open task. The secured check shapes behavior before it even begins. The path of least resistance runs, once again, through understanding the material well enough.

But the evidence here is still thin.

At Wake Forest, only 20 students who hadn’t demonstrated mastery in a written answer were given the chance to take an oral retake. On the cumulative final exam, their scores on the relevant questions caught up with those of students who had handled the topic the first time. It’s an encouraging example of the learning value of oral retrieval, but the study had no proper control group, included restudy and feedback, and used final questions of a familiar type.

In engineering courses at UC San Diego, many students said short oral exams raised their motivation and made them study differently. Yet a small controlled comparison found no convincing improvement in later results. That study is useful precisely for its null result: a feeling of deeper learning is not yet a measured effect.

For now the honest conclusion is this: oral assessment can change how students prepare, retrieve, and explain. Convincing data on any added effect on long-term learning is scarce, and the small studies point in mixed directions.

The same goes for the two-loop model as a whole. Its logic is strong, but direct comparisons with the alternatives are almost nonexistent. We’re seeing the design reform before its pedagogical validation.

/uploads/student-independent-assessment.jpg

Not Oral vs. Written, but Product vs. Verification

The main distinction isn’t between the written and the oral format. It runs between the product and the verification.

Written text remains an important competence. You can’t simply drop essays, reports, and code on the grounds that a machine can now generate them. Graduates still need to research, write, build arguments, and create complex products — now in collaboration with AI as well.

But the finished product can no longer carry the whole evidential load on its own.

The future system will likely assemble a chain of different kinds of evidence: open work with tools, a record of the process, observed practice, a secured checkpoint, delayed recall, and a fresh transfer task. Not every element is needed in every course. What matters is that, by the time the degree is awarded, the university can answer not only “what did the student produce?” but also “what can he understand and do on his own?”

For baseline professional competence, a secured check should perhaps work not as one more slice of a grade average but as a minimum threshold. A flawless AI-assisted project shouldn’t mathematically cancel out the inability to perform a critical action unaided.

This doesn’t mean assigning a universal 30 or 40 percent to the secured part. Research has found no magic weight. The point is a principle: if a program claims a graduate holds a required competence, he should demonstrate it separately, at a minimum acceptable level.

What Will Most Likely Happen Next

The take-home essay probably won’t disappear. It can still teach students to research a topic, form a position, and work alongside AI — but it will gradually lose its status as standalone proof of competence. A strong text will remain a result of work; you just won’t be able to conclude with confidence, from the text alone, who learned what.

The split in assessment is already visible, but not yet universal. Some systems take the redesign route: Sydney builds open and secured loops, Norway trials long-term assignments with a subject conversation, QUT scales oral professional assessment. Others tighten control: they bring back in-person exams, watch screens, use locked-down browsers and network filters. Denmark is moving down both roads at once.

So this isn’t a linear story in which education first banned AI and then wisely embraced it. Control and redesign will coexist. The question is which of them actually measures the competence being claimed, and at what cost.

Educational AI itself will likely grow less obliging. A good tutor will ration help, demand an attempt of your own, return you to material already covered, set transfer problems, and check recall without a prompt. Its quality will have to be judged not by how fast it cleared a difficulty but by what the student can do, later, unaided.

The shape of inequality will change, too. Access will stay important, especially while paid models outperform free ones, but prior knowledge, self-regulation, and the ability to argue with a confident answer will matter more and more. A student with a solid base uses AI as an amplifier; a student without one risks mistaking fluency for truth. Equal access to the model won’t produce equal results.

Finally, a university degree will have to certify, ever more convincingly, the capacity for judgment. Information is already too abundant, answers are becoming nearly free, and the production of outwardly polished results is being automated. Judgment of one’s own — reading the context, spotting the error, owning the choice — stays expensive, and it’s that rare quality universities will have to learn to test.

The Shortest Path Should Run Through Learning

Students are often told to “use AI responsibly.” The advice is right and nearly useless, until someone explains what to actually do at the moment when the full answer is one prompt away.

The practical sequence comes down to five moves: attempt, hint, explain, transfer, reproduce.

First the person solves the task without AI, then asks for minimal help, explains the corrected reasoning in her own words, applies the principle in a different situation, and, after some time, reproduces it unaided. That order keeps the effort that knowledge grows from, while removing the friction that yields nothing.

Universities need a similar discipline. They can assess professional work with AI, but separately obtain secured evidence of minimum independent competence, run delayed checks, set new transfer problems, and publish the results — including the inconvenient ones.

Above all, they have to fix the incentives.

The shortest path to a high grade should run through real learning. Otherwise the ambitious student will, quite rationally, find the way around.

My children will enter this system soon, and I don’t want an AI-free education for them. That world probably won’t exist by then — and the old university hardly deserves unconditional nostalgia anyway. The single pace suited few, late feedback killed curiosity, and routine assignments too often rewarded stamina over understanding.

I want the harder solution. I want teachers to use remarkable tools wisely: to strip out the administrative waste, find the gaps faster, give timely feedback — while keeping the intellectual resistance in the places where thought doesn’t form without it. I want my children challenged, not worn down — and I want the degree to still mean something.

AI really can make education better. But it won’t happen automatically, because markets move faster than institutions, models faster than research, and students adapt fastest of all.

Back to the classroom. Two students earn the top grade, their papers equally impressive, and the teacher could already move on to the next assignment. Instead she asks them both to close their laptops.

“Explain your main decision,” she says. “Now apply it to a different situation.”

Only in that moment can the difference become visible.

The ambitious student didn’t misread the university — on the contrary, he read its grading system a little too well. AI only revealed how little that system can sometimes prove.

The future isn’t a written exam versus an oral one. It’s work with AI plus independent verification — with the method of verification matched to the competence we’re trying to confirm.

If strong student work no longer proves the competence of its author, what does the degree confirm?


Key sources