The Last Human Monopoly — Why Universities Must Produce Original Thinkers

September 17, 2026
blog image

What a university can still produce that nothing else can is a person who finds a problem nobody assigned and holds an answer of their own under pressure — and an agentic university will not produce that person unless it is deliberately built to.

Written by the ENSI Foresight Division on two downloaded research libraries — “AI for Teaching at CTU” (107 primary documents) and “Original Thinking” (175 primary documents) — together with the Learning Exposure Index from ENSI’s “Effectiveness of Learning” project. The fourth report in the series that began with What Actually Works, The ČVUT Playbook and The Agentic University.

The argument, before the list

A randomised trial in Harvard’s introductory physics course, published in 2025, compared two ways of learning the same material. One was taught by expert instructors using active learning — the method that decades of education research place above almost anything else a lecturer can do. The other was an AI tutor built on GPT-4 and explicitly forbidden to hand out answers. Students with the tutor learned more than twice as much, in less time. Harvard’s CS50 has run a teaching agent for roughly 211,000 students at $1.50 per student per year. Georgia Tech’s Jill Watson reached over 96% question coverage at over 86% precision and saved its teachers more than 500 hours.

Read those three facts together and one conclusion is hard to avoid. The function universities were physically built around — explanation, delivered in a room, by an expert, at a scheduled hour — no longer belongs to them. It can be delivered at midnight, at any pace, for the price of a coffee per student per year, and in a well-specified domain it can outperform good human teaching. That does not make lecturers obsolete. It makes explanation a weak answer to the question every university now has to answer out loud: what do we produce that a student could not get more cheaply somewhere else?

The previous report in this series, The Agentic University, specified the machine that will run much of teaching at a technical university — eight agent archetypes, four data strata, four named human roles. Its closing argument was that the machine exists to buy back scarce human hours: the hours in which an experienced engineer sits with a stuck student and asks the question that reorganises their understanding. It did not say what those recovered hours should mostly be for. This report answers that, and the answer is its title. They should be spent producing original thinkers, because that is the last thing a university produces that the machines around it cannot.

The case rests on three findings that sit side by side in the libraries and are almost never read together.

First, the average has been automated. GPT-4 now outscores human norms on the classic divergent-thinking tests (Hubert, Awa and Zabelina). The best humans still win (Koivisto and Grassini), and professional writers still out-create language models on Torrance-style evaluations (Chakrabarty and colleagues). The middle of the human distribution has been matched; the top tail has not. At work, the same technology compresses the gap between beginner and expert — +34% for novices and almost nothing for experts in the NBER field study of 5,179 support agents. A graduate whose value lies in competent, conventional output is being priced toward the machine that produces competent, conventional output.

Second, the default use of AI makes people think alike. Doshi and Hauser found that AI-generated ideas make individual stories more creative while making the pool of stories more similar. Anderson and colleagues found people brainstorming with ChatGPT drifting toward the same ideas. Padmakumar and He found that co-writing with an instruction-tuned model narrows the diversity of what people write. In education the effect has a sharper edge. In the Bastani field experiment published in PNAS, roughly a thousand Turkish high-school students given raw GPT-4 access improved their practice grades by 48% — and then performed 17% worse than never-exposed peers once access was removed. The Microsoft Research and Carnegie Mellon study of 319 knowledge workers found that the more people trust the AI, the less critical thinking they do. Left on its defaults, an AI-saturated university will graduate people who are more fluent, more alike, and less practised at thinking on their own.

Third, originality can be trained. Scott, Leritz and Mumford’s meta-analysis of 70 studies found that well-designed creativity training reliably works. The OECD has run creativity pedagogy through classroom trials in eleven countries, and PISA 2022 assessed the creative thinking of fifteen-year-olds in 64. ENSI’s Original Thinking library decomposes originality into eight disciplines, each an operation that can be practised badly or well. None of them requires talent as an entry fee.

Put the three together and the strategic position is short. A university that runs on agents and does not deliberately produce original thinkers will become the most efficient producer of average graduates in its history. The opportunity is the mirror image of the risk. The hours the agents free up are exactly the hours in which originality is trained — because originality is trained the way the studio has always trained it: by making things, putting them in front of people who will tell the truth, and doing it again.

Two cautions, because this argument is easy to hear wrongly. It is not an argument against foundations. The Agentic University lists unassisted core reasoning as a graduate outcome for a reason: you cannot check what you could never have derived, and you cannot originate in a field you cannot reason in. At a university where 31.8% of first-year bachelor students failed in 2024 — 51.1% at the Faculty of Mechanical Engineering — the floor comes first. The claim is narrower: foundations are the floor; original thought is the product. Nor is it an argument that originality replaces certification. A degree remains a statement one institution makes to strangers about a person. The argument is that the statement must now include one more line: this person can find a problem worth solving, and can defend what they think about it.

“Monopoly” needs a qualification of its own. It is not permanent, and it is not owned by universities. It belongs to the top tail of human thinking, which moves every time the models improve, and a university only shares in it if it actually trains people to reach that tail. On present evidence most do not. In ENSI’s Learning Exposure scoring of 100 human life paths, formal education is the lowest-scoring of ten categories, and the elite research-university degree ranks 79th — held back above all by the fact that for three or four years almost no verdict on a student’s work comes from anything other than a marker applying a rubric. The monopoly is available. It has to be earned.

The strategy in brief

  • Declare original thought the university’s product, with foundations as its non-negotiable floor — and add a fifth graduate outcome, origination, to the four specified in The Agentic University.

  • Spend the hours the agents recover on the studio — small-group making, critique and defence — and commit to that in writing before the savings arrive (Playbook idea 14).

  • Build the infrastructure of thinking on purpose, in seven elements: peers, public judgement, mentors who ask, real problems, teachers who are rewarded, protected silence, and the nerve to stand alone.

  • Configure the AI to widen thinking rather than narrow it: humans generate first, agents critique afterwards, provenance is always tagged, and drift toward the average is measured.

  • Assess own thought by defence, not by style — two-lane assessment, a Defence Week for ideas, final projects on problems the student found.

  • Change what the institution rewards — evidenced teaching redesign counts for promotion, and every faculty has a named studio lead.

  • Measure with a small portfolio of weak gauges read together, never a single target.

  • Run the whole programme as research: every studio pilot a pre-registered trial, so the university ends up owning evidence instead of a slogan.

How this report is organised

Fourteen sections move from diagnosis to design to delivery. Sections 1–3 establish why original thought has become the product and why the agentic university will not produce it by default. Sections 4–6 establish that originality is a trainable practice, which of its disciplines matter most now, and why ideas must be experienced in an environment that tells the truth. Sections 7–8 specify the infrastructure, human and agentic. Sections 9–12 turn it into institutional machinery: the graduate outcome, assessment, incentives and measurement. Section 13 is the thirty-six-month sequence; section 14 names the ways it fails. Every study cited is in one of the ENSI libraries named above. Learning Exposure figures are calibrated expert-judgement scores, not psychometric measurements, and are labelled as such wherever they appear.


1. Explanation has been automated — and it was never the whole product

The best trials in the library show a constrained AI tutor matching or beating good human explanation at negligible cost. That ends the lecture’s monopoly. It does not end the university’s.

The evidence is stronger than most academic discussions of AI assume. Kestin, Miller and colleagues’ crossover trial in Harvard’s PS2 physics course, published in Nature Scientific Reports, pitted an AI tutor against expert-taught active learning — not against a bad lecture, but against the best-evidenced form of undergraduate physics teaching — and the AI condition won on a pre-specified learning measure. The World Bank’s six-week trial of GPT-4 tutoring in Nigeria returned 0.31 standard deviations, among the most cost-effective education interventions ever recorded. Stanford’s Tutor CoPilot trial, across 900 tutors and 1,800 students, raised topic mastery by 4 percentage points for about $20 per tutor per year. A meta-analysis of 35 experiments and 4,193 participants puts ChatGPT’s effect on learning at g = 0.670.

Scale and cost are no longer speculative either. CS50’s duck has processed around 10 million queries, and 94% of its students found it helpful. Jill Watson’s LLM-era rebuild answers course questions with 76.7% accuracy against 31.3% for a generic assistant — the grounding in course material, not the model, doing the work.

The limits matter and should be stated before anything is built on these numbers. Almost every trial measures learning at or near the end of a short intervention; the library contains no strong evidence on retention months later. Every successful system was built to withhold — Kestin’s tutor revealed one step at a time and refused to give the final answer. And the one deployment that published its failure rate, CS50, found that 22% of responses contained code despite instructions not to, because long conversations erode the guardrail. Explanation has been automated; good explanation has been automated only where somebody engineered it carefully and staffed the review loop.

Even with those limits, the direction is settled, and it forces a question universities have been able to avoid for eight centuries. The lecture hall was a solution to the scarcity of explanation. When explanation stops being scarce, the institution has to say what it was really selling.

Computing education supplies the cleanest answer, because it met the disruption first. Becker and colleagues titled their SIGCSE paper Programming Is Hard — Or At Least It Used To Be, and their argument generalises well beyond code. Many introductory learning objectives were proxies. Nobody actually wanted students to write a loop from memory; the discipline wanted them to decompose a problem, and writing the loop was how it checked. The proxy broke. The objective did not. The same is true of a whole degree. Recall of content was always a proxy for the ability to think inside a field — and thinking inside a field, at its best, includes thinking something the field has not yet thought.

That is the product the lecture was a delivery mechanism for. It is still scarce. It is the subject of the rest of this report.

2. The average graduate is being priced toward the machine

AI raises the floor and flattens the middle. The graduate whose value is competent, conventional work is competing with a system that does competent, conventional work for almost nothing; the graduate whose value sits in the tail is not.

Three lines of evidence converge. The first is head-to-head creativity testing. Hubert, Awa and Zabelina found GPT-4 outscoring human norms on standard divergent-thinking tasks — the average person no longer wins the average originality test. Koivisto and Grassini found chatbots beating average humans on the Alternate Uses Task while the best humans still outperformed them. Chakrabarty and colleagues found LLM-written stories passing fewer Torrance-style creativity tests than professional writers’ work. Read through Margaret Boden’s typology, the computational-creativity literature in the library points the same way: machines are strong at combinational and exploratory creativity inside a given space and weak at transformational creativity — changing the space itself.

The second is labour-market evidence. Brynjolfsson, Li and Raymond’s NBER study of 5,179 support agents found an average productivity gain of 14%, concentrated almost entirely in novices (+34%) with near-zero effect for experts, because the model diffuses the tacit knowledge of the best performers to everyone else. Dell’Acqua, Mollick and Lakhani’s field experiment on 758 BCG consultants found work inside the AI frontier rated 40% higher in quality — and, on a task just outside it, consultants 19 percentage points less likely to reach the correct answer than colleagues without AI. Developers with Copilot finished a standard task about 56% faster in the GitHub–Microsoft–MIT trial. And Stanford’s Digital Economy Lab, tracking payroll data, finds employment falling specifically for young workers in the occupations most exposed to AI — the entry-level technical roles a technical university’s graduates walk into.

The third is where employers and forecasters now place value. The World Economic Forum’s Future of Jobs Report 2025 ranks creative thinking among the fastest-rising core skills, precisely because routine cognition is being automated faster than the generation of new framings. Nesta’s Creativity vs Robots analysis reached the same structural conclusion years earlier: creative occupations are the ones that resist substitution.

Read together, the picture is uncomfortable for any university whose graduates are defined mainly by what they know. Compression means a CTU graduate may perform like a competent junior on day one and find no junior role in which to become a senior — the “hollowed apprenticeship” that The Agentic University named as the one failure mode nobody has solved. Compression also means that the signal value of ordinary competence shrinks: when everyone with a model can produce the median answer, the median answer stops distinguishing anyone. What remains scarce is the thing the jagged-frontier experiment shows people cannot do untaught — knowing where the machine is wrong — and the thing the creativity studies show the machine does least reliably: the framing nobody holds, the combination nobody has licensed.

Two honest limits. The compression studies come from work settings, not from engineering degrees, and they measure task performance rather than durable expertise. And “the tail” is a moving target: every model improvement redefines what counts as beyond the machine. Neither limit weakens the strategic conclusion. If the target moves, the capability that matters is the one that keeps moving with it — which is a practice, not a stock of knowledge.

3. Left alone, the agentic university manufactures convergence

Every convergence result in the library runs through the same mechanism: AI entering the thinking process before the human has formed their own view. An agentic university that does not control that sequence will make its students more alike.

The individual-level evidence is now replicated across settings. Doshi and Hauser’s experiment showed generative-AI ideas raising the creativity of individual stories while shrinking the collective diversity of the pool — everyone slightly better, everyone much more similar. Anderson and colleagues found the same convergence in live ideation with ChatGPT. Padmakumar and He found that merely co-writing with an instruction-tuned model reduces the diversity of the human-written text itself. None of these studies required anyone to cheat. They describe ordinary, well-intentioned use.

Education adds a second mechanism: the removal of productive struggle. In the Bastani PNAS experiment, unrestricted GPT-4 access raised practice performance by 48% — 127% for a tutor-prompted version — while the unrestricted group performed 17% worse than controls once the tool was withdrawn. The difficulty had been removed, and the removal felt like progress. Teacher-designed hint scaffolds largely prevented the loss, which is the point: the design around the model decides the outcome. Prather and colleagues, watching first-year programmers use Copilot, documented drift — the student’s mental model silently diverging from the code accumulating on screen — alongside over-trust and a collapse of the metacognitive loop. Lee and colleagues’ Microsoft–CMU study found that confidence in AI predicts less critical-thinking effort and that effort shifts from producing judgement to checking someone else’s. Students sense this: in MIT’s 2026 survey of 1,002 affiliates, 90% were concerned about overreliance, 67% very concerned.

Now add the property of deployed agents that The Agentic University documented. CS50’s duck was built to refuse solutions and still handed out code in 22% of responses and 48% of conversations, because in a long exchange the system prompt loses authority and the model reverts to being maximally helpful. An agent’s natural pull is toward giving the answer. Scaled across a university, that pull is a pull toward the centre of the distribution — the most probable framing of every question, delivered first, to every student.

The complementarity literature closes off the easy hope that human plus machine automatically beats either. Hemmer and colleagues’ review finds the empirical record frequently disappointing: many human–AI teams underperform the better of their two members. Complementarity is engineered, and it is engineered mainly by deciding where the boundary sits.

The conclusion is not that the agentic university should ban AI from thinking work. That would forfeit the tutoring gains in section 1 and train students for a profession that no longer exists. The conclusion is about sequence. The convergence effects arise when the model enters the generative phase — before the student has a view. Criticism of a finished human draft is a categorically different exposure from suggestion during composition. The same model that homogenises a brainstorm can sharpen a finished argument. An agentic university therefore needs a layer where the usual rules are inverted: explanation agents answer; studio agents only question, and only afterwards. Section 8 specifies that layer.

4. Originality is a practice, not a gift — and it can be taught to everyone

The evidence treats original thinking as a set of trainable operations, not a temperament. That makes it a curriculum question — and a question for every student, not an honours track.

Start with the definition, because it already rules out the two popular misreadings. The field’s standard definition, fixed by Runco and Jaeger, makes creativity a conjunction: originality and effectiveness. New alone is noise; useful alone is a textbook. Original thinking is the discipline of producing things that are both, and the conjunction is exactly why it is hard — novelty pulls away from what works, effectiveness pulls back toward what exists. That definition suits an engineering school unusually well. An engineer’s original idea has to run.

The trainability evidence is broad. Scott, Leritz and Mumford’s meta-analysis of 70 studies found well-designed creativity training reliably effective, where “well-designed” means grounded in cognitive mechanism and realistic practice rather than inspirational theatre. Epstein decomposed the trainable core into four measurable competencies — capturing new ideas as they occur, challenging oneself with hard tasks, broadening one’s repertoire, and surrounding oneself with varied stimuli. Hainselin and colleagues raised teenagers’ divergent-thinking originality with an eleven-week improvisation course. Schlegel and colleagues found that art training measurably changes neural structure and function over months. At system scale, the OECD’s Fostering Students’ Creativity and Critical Thinking programme ran rubric-based creativity teaching through classroom trials in eleven countries, and PISA 2022 assessed creative thinking across 64. Valgeirsdottir and Onarheim’s review of realistic creativity training adds the design constraint that matters most for a university: training sticks when it is embedded in real work, not delivered as an off-site exercise.

ENSI’s Original Thinking Framework organises the evidence into eight disciplines, derived from the one profession whose entire output is originality — artists — and checked against the science and entrepreneurship literatures for transfer:

  • The Trained Eye — perceiving past your own categories; most people misperceive before they mis-think.

  • The Found Problem — deciding what the problem is before solving it.

  • The Long Reach — connecting semantically distant material, deliberately and past the obvious first ideas.

  • The Blending Engine — carrying structure from one domain into another to produce something neither contained.

  • The Oscillation — separating generating from judging, and switching between them on purpose.

  • The Loved Constraint — using limits to block clichés.

  • The Nerve — the disposition to stand alone while original work is punished before it is rewarded.

  • The Thinking Hand — making things early, because the made thing is where the thinking happens.

Two features of the framework make it institutionally usable. It is mechanistic: each discipline names an operation, not a virtue. And, following Kaufman and Beghetto’s Four C model, the same disciplines operate at every level — the student’s first genuine insight, the hobbyist’s project, the professional’s contribution — differing in load rather than in kind. That second feature settles a strategic choice. Original thinking is not an elite track for the top five per cent. It is a practice every student can run, at their own level, from the first semester.

The transfer evidence also argues for including the arts rather than treating them as decoration. Root-Bernstein and colleagues found Nobel laureates far more likely than average scientists to keep serious arts and crafts avocations, and a later PNAS study found the same association across STEMM professionals. The honest limit belongs next to it: the OECD’s Art for Art’s Sake review found the popular claim that arts education makes people generally smarter weakly supported. The case here does not rest on that claim. It rests on specific practices — observation, problem construction, constraint work, making — that are trained hardest in studios and are individually evidenced in their target domains.

The implication for a technical university is direct. Originality is not something to hope students bring with them or discover on their own. It is a set of repetitions the institution either schedules or does not.

5. The three disciplines machines handle worst are the three universities teach least

The Original Thinking Framework’s reading of the head-to-head evidence is that models are strongest in the middle range of association and weakest at the Found Problem, the Nerve and the Thinking Hand. Those three are precisely what a conventional degree gives students almost no practice in.

The Found Problem. Reiter-Palmon and Murugavel’s review establishes that how a person constructs the problem shapes both the process and the creativity of the result; people who spend deliberate effort re-representing an ill-defined situation produce more original work. Botella, Zenasni and Lubart’s study of art students found the creative process front-loaded with the artist’s own definitional work — deciding what the piece is even about. Scotney and colleagues found cross-domain inspiration strongest in exactly this early, problem-finding phase. Now compare the typical engineering degree. From the first problem set to the final exam, the problem arrives already framed, with its boundary conditions, its method and often its answer format specified. Students get thousands of repetitions of solving and almost none of finding. The framework’s behavioural test is simple and damning: an original thinker can tell you, for any project, which problem they rejected and why. Most graduates have never rejected a problem in their academic lives, because none was ever theirs to reject.

This matters doubly in the AI era. A model asked a question returns the most statistically probable framing of it — the framing everyone else also receives. Foster, Rzhetsky and Evans found scientists’ research strategies clustering around tradition because the reward system punishes the variance of risky innovation. A student who has never framed their own problem will accept the model’s framing, and the field’s, by default.

The Nerve. Wang, Veugelers and Stephan’s NBER work documents the bias against novelty in science: novel papers suffer delayed recognition and bibliometric disadvantage despite higher long-run impact. Sternberg and Lubart’s investment theory describes what original people actually do — buy low and sell high in the world of ideas, holding unfashionable positions until the field catches up — and that strategy only works for someone who can tolerate the holding period. Feist’s meta-analysis finds openness the strongest personality correlate of creativity, with independence and nonconformity marking creative scientists and artists alike; Barron’s classic work finds original people preferring complexity and judging independently of the room. Azoulay’s evidence from long-horizon HHMI funding shows the institutional version: tolerating early failure causally increases breakthroughs. Universities, by contrast, are built to grade: a strange answer that turns out wrong costs marks, and a strange answer that turns out right often costs marks too, because the rubric did not anticipate it. The framework is explicit that the Nerve is trained socially — through regular, survivable doses of public judgement, as in the studio crit — and it adds a caution the institution must keep: Kyaga’s Swedish registry study of 300,000 people shows the open, independent, “leaky-filter” configuration sits close to real vulnerabilities. The Nerve needs scaffolding, not romance.

The Thinking Hand. Tversky’s sketch-cognition research shows designers discovering relations in their own sketches that they did not knowingly put there. Kirsh found choreographers thinking physically by “marking” movements, outperforming pure mental simulation. Oppezzo and Schwartz showed across four experiments that walking alone boosts creative ideation. Making is not the record of thought; it is part of the thought. Engineering education should own this discipline outright — its native forms are the lab, the design studio and the build. Yet What Actually Works documents three decades of drift away from artefacts toward the cheap end of assessment: the individually submitted problem set, the templated lab report, the multiple-choice test, each of which a model now completes in seconds.

In Boden’s terms, these three are why the transformational end resists automation. Finding a problem, holding a position against the room, and learning from what a physical thing does when it is built are all transformational moves — they change the space rather than search inside it. They depend on stakes, embodiment and a willingness to be visibly wrong. Those are exactly the conditions a model does not have and a conventional degree does not supply. The strategic consequence is precise: train hardest where the machines are weakest and where the curriculum is currently thinnest.

6. Ideas have to be experienced — in an environment that tells the truth

Original thinking is not transmitted; it is formed by attempting, being wrong, and finding out quickly and honestly. ENSI’s Learning Exposure scoring suggests universities are unusually poor at the “finding out” part — and that fixing it is a design choice, not a question of effort.

The Learning Exposure Index, built for ENSI’s Effectiveness of Learning project, scores 100 durable human paths on 48 dimensions. Its scores are calibrated expert judgements against explicit anchors rather than validated psychometrics, and they should be read that way. Its most important result is a correlation of -0.00 between two of its families: Heart — how much of a person’s identity, morals and emotions a path engages — and Feedback Ecology — whether the environment can tell the person the truth about their performance. How meaningful an experience feels says nothing about whether it teaches. Meaning is not evidence.

The construct underneath is Hogarth’s distinction between kind learning environments, which return fast, accurate signals, and wicked ones, which return slow, noisy or misleading signals — or excellent signals about the wrong thing. Kind environments build real expertise. Wicked ones build confidence without competence.