

What a university can still produce that nothing else can is a person who finds a problem nobody assigned and holds an answer of their own under pressure — and an agentic university will not produce that person unless it is deliberately built to.
Written by the ENSI Foresight Division on two downloaded research libraries — “AI for Teaching at CTU” (107 primary documents) and “Original Thinking” (175 primary documents) — together with the Learning Exposure Index from ENSI’s “Effectiveness of Learning” project. The fourth report in the series that began with What Actually Works, The ČVUT Playbook and The Agentic University.
A randomised trial in Harvard’s introductory physics course, published in 2025, compared two ways of learning the same material. One was taught by expert instructors using active learning — the method that decades of education research place above almost anything else a lecturer can do. The other was an AI tutor built on GPT-4 and explicitly forbidden to hand out answers. Students with the tutor learned more than twice as much, in less time. Harvard’s CS50 has run a teaching agent for roughly 211,000 students at $1.50 per student per year. Georgia Tech’s Jill Watson reached over 96% question coverage at over 86% precision and saved its teachers more than 500 hours.
Read those three facts together and one conclusion is hard to avoid. The function universities were physically built around — explanation, delivered in a room, by an expert, at a scheduled hour — no longer belongs to them. It can be delivered at midnight, at any pace, for the price of a coffee per student per year, and in a well-specified domain it can outperform good human teaching. That does not make lecturers obsolete. It makes explanation a weak answer to the question every university now has to answer out loud: what do we produce that a student could not get more cheaply somewhere else?
The previous report in this series, The Agentic University, specified the machine that will run much of teaching at a technical university — eight agent archetypes, four data strata, four named human roles. Its closing argument was that the machine exists to buy back scarce human hours: the hours in which an experienced engineer sits with a stuck student and asks the question that reorganises their understanding. It did not say what those recovered hours should mostly be for. This report answers that, and the answer is its title. They should be spent producing original thinkers, because that is the last thing a university produces that the machines around it cannot.
The case rests on three findings that sit side by side in the libraries and are almost never read together.
First, the average has been automated. GPT-4 now outscores human norms on the classic divergent-thinking tests (Hubert, Awa and Zabelina). The best humans still win (Koivisto and Grassini), and professional writers still out-create language models on Torrance-style evaluations (Chakrabarty and colleagues). The middle of the human distribution has been matched; the top tail has not. At work, the same technology compresses the gap between beginner and expert — +34% for novices and almost nothing for experts in the NBER field study of 5,179 support agents. A graduate whose value lies in competent, conventional output is being priced toward the machine that produces competent, conventional output.
Second, the default use of AI makes people think alike. Doshi and Hauser found that AI-generated ideas make individual stories more creative while making the pool of stories more similar. Anderson and colleagues found people brainstorming with ChatGPT drifting toward the same ideas. Padmakumar and He found that co-writing with an instruction-tuned model narrows the diversity of what people write. In education the effect has a sharper edge. In the Bastani field experiment published in PNAS, roughly a thousand Turkish high-school students given raw GPT-4 access improved their practice grades by 48% — and then performed 17% worse than never-exposed peers once access was removed. The Microsoft Research and Carnegie Mellon study of 319 knowledge workers found that the more people trust the AI, the less critical thinking they do. Left on its defaults, an AI-saturated university will graduate people who are more fluent, more alike, and less practised at thinking on their own.
Third, originality can be trained. Scott, Leritz and Mumford’s meta-analysis of 70 studies found that well-designed creativity training reliably works. The OECD has run creativity pedagogy through classroom trials in eleven countries, and PISA 2022 assessed the creative thinking of fifteen-year-olds in 64. ENSI’s Original Thinking library decomposes originality into eight disciplines, each an operation that can be practised badly or well. None of them requires talent as an entry fee.
Put the three together and the strategic position is short. A university that runs on agents and does not deliberately produce original thinkers will become the most efficient producer of average graduates in its history. The opportunity is the mirror image of the risk. The hours the agents free up are exactly the hours in which originality is trained — because originality is trained the way the studio has always trained it: by making things, putting them in front of people who will tell the truth, and doing it again.
Two cautions, because this argument is easy to hear wrongly. It is not an argument against foundations. The Agentic University lists unassisted core reasoning as a graduate outcome for a reason: you cannot check what you could never have derived, and you cannot originate in a field you cannot reason in. At a university where 31.8% of first-year bachelor students failed in 2024 — 51.1% at the Faculty of Mechanical Engineering — the floor comes first. The claim is narrower: foundations are the floor; original thought is the product. Nor is it an argument that originality replaces certification. A degree remains a statement one institution makes to strangers about a person. The argument is that the statement must now include one more line: this person can find a problem worth solving, and can defend what they think about it.
“Monopoly” needs a qualification of its own. It is not permanent, and it is not owned by universities. It belongs to the top tail of human thinking, which moves every time the models improve, and a university only shares in it if it actually trains people to reach that tail. On present evidence most do not. In ENSI’s Learning Exposure scoring of 100 human life paths, formal education is the lowest-scoring of ten categories, and the elite research-university degree ranks 79th — held back above all by the fact that for three or four years almost no verdict on a student’s work comes from anything other than a marker applying a rubric. The monopoly is available. It has to be earned.
Declare original thought the university’s product, with foundations as its non-negotiable floor — and add a fifth graduate outcome, origination, to the four specified in The Agentic University.
Spend the hours the agents recover on the studio — small-group making, critique and defence — and commit to that in writing before the savings arrive (Playbook idea 14).
Build the infrastructure of thinking on purpose, in seven elements: peers, public judgement, mentors who ask, real problems, teachers who are rewarded, protected silence, and the nerve to stand alone.
Configure the AI to widen thinking rather than narrow it: humans generate first, agents critique afterwards, provenance is always tagged, and drift toward the average is measured.
Assess own thought by defence, not by style — two-lane assessment, a Defence Week for ideas, final projects on problems the student found.
Change what the institution rewards — evidenced teaching redesign counts for promotion, and every faculty has a named studio lead.
Measure with a small portfolio of weak gauges read together, never a single target.
Run the whole programme as research: every studio pilot a pre-registered trial, so the university ends up owning evidence instead of a slogan.
Fourteen sections move from diagnosis to design to delivery. Sections 1–3 establish why original thought has become the product and why the agentic university will not produce it by default. Sections 4–6 establish that originality is a trainable practice, which of its disciplines matter most now, and why ideas must be experienced in an environment that tells the truth. Sections 7–8 specify the infrastructure, human and agentic. Sections 9–12 turn it into institutional machinery: the graduate outcome, assessment, incentives and measurement. Section 13 is the thirty-six-month sequence; section 14 names the ways it fails. Every study cited is in one of the ENSI libraries named above. Learning Exposure figures are calibrated expert-judgement scores, not psychometric measurements, and are labelled as such wherever they appear.
The best trials in the library show a constrained AI tutor matching or beating good human explanation at negligible cost. That ends the lecture’s monopoly. It does not end the university’s.
The evidence is stronger than most academic discussions of AI assume. Kestin, Miller and colleagues’ crossover trial in Harvard’s PS2 physics course, published in Nature Scientific Reports, pitted an AI tutor against expert-taught active learning — not against a bad lecture, but against the best-evidenced form of undergraduate physics teaching — and the AI condition won on a pre-specified learning measure. The World Bank’s six-week trial of GPT-4 tutoring in Nigeria returned 0.31 standard deviations, among the most cost-effective education interventions ever recorded. Stanford’s Tutor CoPilot trial, across 900 tutors and 1,800 students, raised topic mastery by 4 percentage points for about $20 per tutor per year. A meta-analysis of 35 experiments and 4,193 participants puts ChatGPT’s effect on learning at g = 0.670.
Scale and cost are no longer speculative either. CS50’s duck has processed around 10 million queries, and 94% of its students found it helpful. Jill Watson’s LLM-era rebuild answers course questions with 76.7% accuracy against 31.3% for a generic assistant — the grounding in course material, not the model, doing the work.
The limits matter and should be stated before anything is built on these numbers. Almost every trial measures learning at or near the end of a short intervention; the library contains no strong evidence on retention months later. Every successful system was built to withhold — Kestin’s tutor revealed one step at a time and refused to give the final answer. And the one deployment that published its failure rate, CS50, found that 22% of responses contained code despite instructions not to, because long conversations erode the guardrail. Explanation has been automated; good explanation has been automated only where somebody engineered it carefully and staffed the review loop.
Even with those limits, the direction is settled, and it forces a question universities have been able to avoid for eight centuries. The lecture hall was a solution to the scarcity of explanation. When explanation stops being scarce, the institution has to say what it was really selling.
Computing education supplies the cleanest answer, because it met the disruption first. Becker and colleagues titled their SIGCSE paper Programming Is Hard — Or At Least It Used To Be, and their argument generalises well beyond code. Many introductory learning objectives were proxies. Nobody actually wanted students to write a loop from memory; the discipline wanted them to decompose a problem, and writing the loop was how it checked. The proxy broke. The objective did not. The same is true of a whole degree. Recall of content was always a proxy for the ability to think inside a field — and thinking inside a field, at its best, includes thinking something the field has not yet thought.
That is the product the lecture was a delivery mechanism for. It is still scarce. It is the subject of the rest of this report.
AI raises the floor and flattens the middle. The graduate whose value is competent, conventional work is competing with a system that does competent, conventional work for almost nothing; the graduate whose value sits in the tail is not.
Three lines of evidence converge. The first is head-to-head creativity testing. Hubert, Awa and Zabelina found GPT-4 outscoring human norms on standard divergent-thinking tasks — the average person no longer wins the average originality test. Koivisto and Grassini found chatbots beating average humans on the Alternate Uses Task while the best humans still outperformed them. Chakrabarty and colleagues found LLM-written stories passing fewer Torrance-style creativity tests than professional writers’ work. Read through Margaret Boden’s typology, the computational-creativity literature in the library points the same way: machines are strong at combinational and exploratory creativity inside a given space and weak at transformational creativity — changing the space itself.
The second is labour-market evidence. Brynjolfsson, Li and Raymond’s NBER study of 5,179 support agents found an average productivity gain of 14%, concentrated almost entirely in novices (+34%) with near-zero effect for experts, because the model diffuses the tacit knowledge of the best performers to everyone else. Dell’Acqua, Mollick and Lakhani’s field experiment on 758 BCG consultants found work inside the AI frontier rated 40% higher in quality — and, on a task just outside it, consultants 19 percentage points less likely to reach the correct answer than colleagues without AI. Developers with Copilot finished a standard task about 56% faster in the GitHub–Microsoft–MIT trial. And Stanford’s Digital Economy Lab, tracking payroll data, finds employment falling specifically for young workers in the occupations most exposed to AI — the entry-level technical roles a technical university’s graduates walk into.
The third is where employers and forecasters now place value. The World Economic Forum’s Future of Jobs Report 2025 ranks creative thinking among the fastest-rising core skills, precisely because routine cognition is being automated faster than the generation of new framings. Nesta’s Creativity vs Robots analysis reached the same structural conclusion years earlier: creative occupations are the ones that resist substitution.
Read together, the picture is uncomfortable for any university whose graduates are defined mainly by what they know. Compression means a CTU graduate may perform like a competent junior on day one and find no junior role in which to become a senior — the “hollowed apprenticeship” that The Agentic University named as the one failure mode nobody has solved. Compression also means that the signal value of ordinary competence shrinks: when everyone with a model can produce the median answer, the median answer stops distinguishing anyone. What remains scarce is the thing the jagged-frontier experiment shows people cannot do untaught — knowing where the machine is wrong — and the thing the creativity studies show the machine does least reliably: the framing nobody holds, the combination nobody has licensed.
Two honest limits. The compression studies come from work settings, not from engineering degrees, and they measure task performance rather than durable expertise. And “the tail” is a moving target: every model improvement redefines what counts as beyond the machine. Neither limit weakens the strategic conclusion. If the target moves, the capability that matters is the one that keeps moving with it — which is a practice, not a stock of knowledge.
Every convergence result in the library runs through the same mechanism: AI entering the thinking process before the human has formed their own view. An agentic university that does not control that sequence will make its students more alike.
The individual-level evidence is now replicated across settings. Doshi and Hauser’s experiment showed generative-AI ideas raising the creativity of individual stories while shrinking the collective diversity of the pool — everyone slightly better, everyone much more similar. Anderson and colleagues found the same convergence in live ideation with ChatGPT. Padmakumar and He found that merely co-writing with an instruction-tuned model reduces the diversity of the human-written text itself. None of these studies required anyone to cheat. They describe ordinary, well-intentioned use.
Education adds a second mechanism: the removal of productive struggle. In the Bastani PNAS experiment, unrestricted GPT-4 access raised practice performance by 48% — 127% for a tutor-prompted version — while the unrestricted group performed 17% worse than controls once the tool was withdrawn. The difficulty had been removed, and the removal felt like progress. Teacher-designed hint scaffolds largely prevented the loss, which is the point: the design around the model decides the outcome. Prather and colleagues, watching first-year programmers use Copilot, documented drift — the student’s mental model silently diverging from the code accumulating on screen — alongside over-trust and a collapse of the metacognitive loop. Lee and colleagues’ Microsoft–CMU study found that confidence in AI predicts less critical-thinking effort and that effort shifts from producing judgement to checking someone else’s. Students sense this: in MIT’s 2026 survey of 1,002 affiliates, 90% were concerned about overreliance, 67% very concerned.
Now add the property of deployed agents that The Agentic University documented. CS50’s duck was built to refuse solutions and still handed out code in 22% of responses and 48% of conversations, because in a long exchange the system prompt loses authority and the model reverts to being maximally helpful. An agent’s natural pull is toward giving the answer. Scaled across a university, that pull is a pull toward the centre of the distribution — the most probable framing of every question, delivered first, to every student.
The complementarity literature closes off the easy hope that human plus machine automatically beats either. Hemmer and colleagues’ review finds the empirical record frequently disappointing: many human–AI teams underperform the better of their two members. Complementarity is engineered, and it is engineered mainly by deciding where the boundary sits.
The conclusion is not that the agentic university should ban AI from thinking work. That would forfeit the tutoring gains in section 1 and train students for a profession that no longer exists. The conclusion is about sequence. The convergence effects arise when the model enters the generative phase — before the student has a view. Criticism of a finished human draft is a categorically different exposure from suggestion during composition. The same model that homogenises a brainstorm can sharpen a finished argument. An agentic university therefore needs a layer where the usual rules are inverted: explanation agents answer; studio agents only question, and only afterwards. Section 8 specifies that layer.
The evidence treats original thinking as a set of trainable operations, not a temperament. That makes it a curriculum question — and a question for every student, not an honours track.
Start with the definition, because it already rules out the two popular misreadings. The field’s standard definition, fixed by Runco and Jaeger, makes creativity a conjunction: originality and effectiveness. New alone is noise; useful alone is a textbook. Original thinking is the discipline of producing things that are both, and the conjunction is exactly why it is hard — novelty pulls away from what works, effectiveness pulls back toward what exists. That definition suits an engineering school unusually well. An engineer’s original idea has to run.
The trainability evidence is broad. Scott, Leritz and Mumford’s meta-analysis of 70 studies found well-designed creativity training reliably effective, where “well-designed” means grounded in cognitive mechanism and realistic practice rather than inspirational theatre. Epstein decomposed the trainable core into four measurable competencies — capturing new ideas as they occur, challenging oneself with hard tasks, broadening one’s repertoire, and surrounding oneself with varied stimuli. Hainselin and colleagues raised teenagers’ divergent-thinking originality with an eleven-week improvisation course. Schlegel and colleagues found that art training measurably changes neural structure and function over months. At system scale, the OECD’s Fostering Students’ Creativity and Critical Thinking programme ran rubric-based creativity teaching through classroom trials in eleven countries, and PISA 2022 assessed creative thinking across 64. Valgeirsdottir and Onarheim’s review of realistic creativity training adds the design constraint that matters most for a university: training sticks when it is embedded in real work, not delivered as an off-site exercise.
ENSI’s Original Thinking Framework organises the evidence into eight disciplines, derived from the one profession whose entire output is originality — artists — and checked against the science and entrepreneurship literatures for transfer:
The Trained Eye — perceiving past your own categories; most people misperceive before they mis-think.
The Found Problem — deciding what the problem is before solving it.
The Long Reach — connecting semantically distant material, deliberately and past the obvious first ideas.
The Blending Engine — carrying structure from one domain into another to produce something neither contained.
The Oscillation — separating generating from judging, and switching between them on purpose.
The Loved Constraint — using limits to block clichés.
The Nerve — the disposition to stand alone while original work is punished before it is rewarded.
The Thinking Hand — making things early, because the made thing is where the thinking happens.
Two features of the framework make it institutionally usable. It is mechanistic: each discipline names an operation, not a virtue. And, following Kaufman and Beghetto’s Four C model, the same disciplines operate at every level — the student’s first genuine insight, the hobbyist’s project, the professional’s contribution — differing in load rather than in kind. That second feature settles a strategic choice. Original thinking is not an elite track for the top five per cent. It is a practice every student can run, at their own level, from the first semester.
The transfer evidence also argues for including the arts rather than treating them as decoration. Root-Bernstein and colleagues found Nobel laureates far more likely than average scientists to keep serious arts and crafts avocations, and a later PNAS study found the same association across STEMM professionals. The honest limit belongs next to it: the OECD’s Art for Art’s Sake review found the popular claim that arts education makes people generally smarter weakly supported. The case here does not rest on that claim. It rests on specific practices — observation, problem construction, constraint work, making — that are trained hardest in studios and are individually evidenced in their target domains.
The implication for a technical university is direct. Originality is not something to hope students bring with them or discover on their own. It is a set of repetitions the institution either schedules or does not.
The Original Thinking Framework’s reading of the head-to-head evidence is that models are strongest in the middle range of association and weakest at the Found Problem, the Nerve and the Thinking Hand. Those three are precisely what a conventional degree gives students almost no practice in.
The Found Problem. Reiter-Palmon and Murugavel’s review establishes that how a person constructs the problem shapes both the process and the creativity of the result; people who spend deliberate effort re-representing an ill-defined situation produce more original work. Botella, Zenasni and Lubart’s study of art students found the creative process front-loaded with the artist’s own definitional work — deciding what the piece is even about. Scotney and colleagues found cross-domain inspiration strongest in exactly this early, problem-finding phase. Now compare the typical engineering degree. From the first problem set to the final exam, the problem arrives already framed, with its boundary conditions, its method and often its answer format specified. Students get thousands of repetitions of solving and almost none of finding. The framework’s behavioural test is simple and damning: an original thinker can tell you, for any project, which problem they rejected and why. Most graduates have never rejected a problem in their academic lives, because none was ever theirs to reject.
This matters doubly in the AI era. A model asked a question returns the most statistically probable framing of it — the framing everyone else also receives. Foster, Rzhetsky and Evans found scientists’ research strategies clustering around tradition because the reward system punishes the variance of risky innovation. A student who has never framed their own problem will accept the model’s framing, and the field’s, by default.
The Nerve. Wang, Veugelers and Stephan’s NBER work documents the bias against novelty in science: novel papers suffer delayed recognition and bibliometric disadvantage despite higher long-run impact. Sternberg and Lubart’s investment theory describes what original people actually do — buy low and sell high in the world of ideas, holding unfashionable positions until the field catches up — and that strategy only works for someone who can tolerate the holding period. Feist’s meta-analysis finds openness the strongest personality correlate of creativity, with independence and nonconformity marking creative scientists and artists alike; Barron’s classic work finds original people preferring complexity and judging independently of the room. Azoulay’s evidence from long-horizon HHMI funding shows the institutional version: tolerating early failure causally increases breakthroughs. Universities, by contrast, are built to grade: a strange answer that turns out wrong costs marks, and a strange answer that turns out right often costs marks too, because the rubric did not anticipate it. The framework is explicit that the Nerve is trained socially — through regular, survivable doses of public judgement, as in the studio crit — and it adds a caution the institution must keep: Kyaga’s Swedish registry study of 300,000 people shows the open, independent, “leaky-filter” configuration sits close to real vulnerabilities. The Nerve needs scaffolding, not romance.
The Thinking Hand. Tversky’s sketch-cognition research shows designers discovering relations in their own sketches that they did not knowingly put there. Kirsh found choreographers thinking physically by “marking” movements, outperforming pure mental simulation. Oppezzo and Schwartz showed across four experiments that walking alone boosts creative ideation. Making is not the record of thought; it is part of the thought. Engineering education should own this discipline outright — its native forms are the lab, the design studio and the build. Yet What Actually Works documents three decades of drift away from artefacts toward the cheap end of assessment: the individually submitted problem set, the templated lab report, the multiple-choice test, each of which a model now completes in seconds.
In Boden’s terms, these three are why the transformational end resists automation. Finding a problem, holding a position against the room, and learning from what a physical thing does when it is built are all transformational moves — they change the space rather than search inside it. They depend on stakes, embodiment and a willingness to be visibly wrong. Those are exactly the conditions a model does not have and a conventional degree does not supply. The strategic consequence is precise: train hardest where the machines are weakest and where the curriculum is currently thinnest.
Original thinking is not transmitted; it is formed by attempting, being wrong, and finding out quickly and honestly. ENSI’s Learning Exposure scoring suggests universities are unusually poor at the “finding out” part — and that fixing it is a design choice, not a question of effort.
The Learning Exposure Index, built for ENSI’s Effectiveness of Learning project, scores 100 durable human paths on 48 dimensions. Its scores are calibrated expert judgements against explicit anchors rather than validated psychometrics, and they should be read that way. Its most important result is a correlation of -0.00 between two of its families: Heart — how much of a person’s identity, morals and emotions a path engages — and Feedback Ecology — whether the environment can tell the person the truth about their performance. How meaningful an experience feels says nothing about whether it teaches. Meaning is not evidence.
The construct underneath is Hogarth’s distinction between kind learning environments, which return fast, accurate signals, and wicked ones, which return slow, noisy or misleading signals — or excellent signals about the wrong thing. Kind environments build real expertise. Wicked ones build confidence without competence.
Scored against that distinction, formal education comes out badly. It is the lowest-scoring of ten categories, at a mean index of 42.3; creative practice is the highest at 62.1. The elite research-university degree ranks 79th of 100 and the mass-market university 93rd. The mechanism is specific. The elite-university dossier scores reality contact at 1: for three or four years, every verdict on a student’s work comes from an intermediary with a rubric — never a user, a client, an opponent or a physical system. Feedback fidelity and latency score 2; mastery legibility 2, meaning a student can finish a famous degree without knowing how good they are. Consequence weight is 1 while identity entanglement is 3 — a high stake in the self sitting on top of a minimal stake in the world.
Independent measurement points the same way, with a necessary counterweight. Arum, Roksa and Cho’s longitudinal study found gains of only about 0.18 standard deviations in critical thinking, complex reasoning and writing over the first two years of college. Mountjoy and Hickman found that selectivity barely predicts a college’s value-added. But Ritchie and Tucker-Drob’s meta-analysis finds each additional year of education raising measured cognitive ability by roughly 1 to 5 IQ points, and the earnings premium is real. The claim is not that university fails to build capability. It is that universities convert a small share of a student’s years into the kind of experience that forms independent judgement — because the verdict almost never comes from reality.
The learning-science literature says what a better environment looks like, and it contains a distinction that is routinely collapsed. Bjork and Bjork’s desirable difficulties — spacing, interleaving, retrieval, generation — make learning feel harder and produce more durable results; Roediger and Karpicke found tested material recalled at 61% against 40% for restudied material after a week. Kapur’s productive-failure studies found students who attempted ill-structured problems before instruction outperforming directly-instructed peers on conceptual understanding and transfer. Metcalfe found errorful generation followed by correction beats errorless study. Every one of these is difficulty with feedback attached. Difficulty without feedback is not desirable; it is opacity. Productive failure is productive only because instruction follows it.
Two more dimensions carry the design. The first is error affordability — whether being wrong is cheap enough to repeat. It is the only one of the index’s 48 dimensions that correlates negatively with total exposure (-0.13): demanding environments tend to make failure expensive. The inversions are instructive. Stand-up comedy scores error affordability at the maximum and feedback fidelity at the maximum — laughter is involuntary and arrives within a second — and ranks joint second of all 100 paths. Commercial aviation decoupled consequence from cost by building the simulator, and McGaghie and colleagues’ synthesis shows the medical equivalent, simulation-based mastery learning, improving real patient outcomes. The second is deliberate-practice affordance — whether the hard part can be isolated and repeated. Macnamara and colleagues’ meta-analysis of 88 studies found practice explaining about 26% of performance variance in games but under 1% in professions, a gradient that tracks how structured the environment is rather than how hard people work. Deans for Impact’s summary is blunt: most experience does not produce expertise unless it is deliberately structured.
There is a warning inside the same data. Studio visual art has the highest “grip” score in the catalogue — the most absorbing, most personally engaging path — alongside a feedback ecology of only 41.7. An environment can hold a person completely for a decade and never once tell them the truth. A university studio built on the art-school model alone would reproduce that failure: intense, meaningful and uninformative. The model to copy is a hybrid — the studio’s making, the comedy club’s fast and honest audience, and the simulator’s cheap failure.
That hybrid is closer to hand at a technical university than anywhere else. What Actually Works calls engineering’s artefact-based pedagogy the structural advantage nobody is exploiting: a structure is checked by statics, a circuit oscillates or it does not, a robot finishes the task or falls over. The verdict is external to the examiner. So the design brief for “experiencing ideas” reduces to four levers the index isolates:
Raise reality contact — let a user, a client, a physical system or a competitor deliver the verdict wherever possible.
Make error cheap — simulators, digital twins, prototypes, rehearsal defences, repeated attempts with no grade attached.
Shorten the loop — feedback in days, not at the end of the semester.
Isolate the hard part — let students repeat problem framing, critique and defence as separate, practised skills.
None of those levers requires more hours from anyone or more rigour. They require the environment to be honest, on a schedule, at a cost the student can afford.
Independent thought does not appear on its own, and it does not come from lectures or well-meant advice. It forms in an environment that pushes back. That environment has seven parts, and at most universities each of them exists only by accident — usually because one teacher builds it in spite of the system.
A school has buildings, curricula and examinations. What it rarely has, by design, is the set of conditions under which a young person starts thinking for themselves and keeps doing it when that becomes costly. The seven elements below are that set. Each is specified the same way: what it does, what the evidence says, what it looks like at CTU, what the agents may do, and the boundary that must hold.
What it does. An idea that never leaves your head is not yet an idea. Equals force it out — they ask what you actually claim, and they build on it or break it.
The evidence. Sawyer and DeZutter’s work on collaborative emergence shows collective originality arising from visible contributions others can build on; nothing can be built on what stays private. Glăveanu’s craft-ecology fieldwork shows creative communities deciding, together, which departures from tradition count as contributions. Dunbar’s in-vivo studies of molecular-biology labs found discoveries emerging from distributed reasoning and analogy in ordinary lab meetings. Wu, Wang and Evans, across 65 million papers, patents and products, found that small teams disrupt and large teams develop.
At CTU. Studio cohorts of roughly eight to twelve students, deliberately mixed across the eight faculties, running from the second year onward. Cross-faculty build projects of the kind the Playbook proposes (idea 20: FIT and FEL students building tutoring agents for FS and FSv courses, with the receiving academic as client).
The agents. The Archivist keeps each cohort’s shared corpus of sketches, drafts and dead ends, with every item tagged by provenance.
The boundary. No agent takes part in a group’s generative session. Groups present rough artefacts early; the polished pitch is banned in the first round, because collective originality needs something unfinished to build on.
What it does. Without comparison against others, a person never discovers where their real limits are — and that they usually sit further out than they thought.
The evidence. The strongest evidence in the libraries concerns exposure rather than competition as such. Stand-up comedy — publicly attributed work, weekly rejection, an audience that cannot fake a laugh — scores at the maximum on feedback fidelity and error affordability and ranks joint second of 100 paths in the Learning Exposure scoring. The studio crit is described in the Original Thinking library as a deliberate exposure regime: regular, survivable doses of public judgement. The same scoring carries a warning: when the verdict comes from a panel applying its taste, the fidelity ceiling is low however sophisticated the panel.
At CTU. Open-problem challenges with industry and public-sector partners, judged wherever possible by something that does not care about the entrant — a benchmark, a live user, a working prototype, an adversarial test. The university’s existing talent environments, such as FIKS at FIT and FEL Camp, are the seeds; the Playbook already proposes scaling them (idea 15).
The agents. The Red-Team forecasts how judges and users will attack an entry, so the entrant walks in prepared rather than protected.
The boundary. Winning must depend on what works, not on polish or presentation. A competition judged on slides trains slide-making.
What it does. A mentor does not supply the conclusion. They show the student how to question their own framing, and they are close enough to say, specifically, “this is wrong”.
The evidence. Every AI tutor that produced large learning gains was built to withhold. Stanford’s Tutor CoPilot did not tutor students at all; it coached human tutors in real time, and helped students of the weakest tutors most (+9 percentage points). A 2025 systematic review of deliberate practice in psychotherapy training found coached, feedback-rich practice outperforming standard training even in open-ended, judgement-heavy work. In the Learning Exposure scoring, mentorship density scores 4 for airline pilots and emergency-medicine residents — and 1 for the mass-market university, whose dossier records “a striking absence of anyone who ever learned their name well enough to tell them they were wrong.”
At CTU. Every student in the studio years has one named mentor who knows their work. Doctoral students at CIIRC and the faculties’ research groups are a natural mentor pool, and the Playbook’s proposal to make teaching-AI a doctoral topic (idea 26) extends to studio practice. Industry engineers mentor through the Junior Engineer Compact (idea 4).
The agents. The Studio Critic — already specified in The Agentic University as “critique, never generation” — questions students’ work between sessions. A Tutor CoPilot–style assistant can coach the human mentor on what to ask.
The boundary. Agents may ask questions; they may not supply a framing. The mentor’s job, human or machine, is to make the student’s thinking better, not to replace it with better thinking.
What it does. An idea tested only against a rubric has never been tested. Real problems return a verdict the student cannot argue with.
The evidence. Reality contact is the dimension on which the elite university scores 1 and the one the Learning Exposure work calls the single highest-leverage change available. Emergency-medicine residency shows high stakes combined with fast, truthful feedback. Sarasvathy’s expert entrepreneurs structure attempts around affordable loss, so being wrong is survivable and therefore repeatable. Camuffo and colleagues’ randomised trial found founders trained to treat beliefs as testable hypotheses made measurably better pivot decisions. CESAER’s Engineer of the Future white paper and the CDIO tradition already commit European technical universities to challenge-based learning.
At CTU. The Junior Engineer Compact — supervised responsibility for real work in the final year — and a pipeline of real problems through EDIH CTU, the AI-MATTERS testing facility and the university’s industrial partners. Strategic-plan goal 1.3, “bring practice into teaching”, is the existing mandate. Crucially, the problem arrives as a situation, not a specification: the student’s first deliverable is the problem statement.
The agents. The Scout brings in mechanisms from distant fields; the Lab and Simulation Agent provides digital twins so that failure is cheap before it is expensive.
The boundary. The partner gets useful work; CTU certifies the judgement. If assessment is of output alone, the scheme becomes free labour, which the Playbook names as an ethical failure mode.
What it does. Every element above is delivered by academics. Studio teaching is harder than lecturing: it requires tolerating ambiguity, critiquing without taking over, and being pleased when a student sees further than the teacher.
The evidence. Ithaka S+R’s interviews across nineteen universities found teaching innovation happening in isolated pockets and dying on friction — no time, no recognition, nobody to ask. The multi-institution barriers study found those obstacles operating independently at individual, departmental and institutional level, so fixing one changes little. The Playbook rates changing promotion criteria as the highest-leverage idea on its list and notes that not one study has tested it.
At CTU. 2,247 academic staff whose careers currently reward publication. Evidenced teaching redesign counts for promotion (idea 6); redesign fellowships give released time with a deliverable (idea 17); ETH Zurich’s lecturer framework is adopted rather than rewritten (idea 11), extended with studio skills such as running a critique.
The agents. The Course Concierge and Instructional Design Agent absorb logistics and first drafts, which is where the teaching hours come from (section 8).
The boundary. Studio teaching is credited, not added. Asking academics to run studios on top of their existing load is how the whole strategy fails quietly.
What it does. Without periods in which nobody wants anything, an independent thought has no moment to surface. In an AI-saturated environment, silence has to be scheduled.
The evidence. Sio and Ormerod’s meta-analysis in Psychological Bulletin found incubation effects real and positive, strongest for divergent problems; Gilhooly’s work supplies the mechanism of unconscious processing. Oppezzo and Schwartz found that walking boosts creative ideation. Ellamil and colleagues showed that generating and evaluating are different brain states, which is why the studio separates them. Torrance built incubation into his model of teaching decades ago. And the Bastani experiment shows what the absence of struggle costs.
At CTU. Projects are opened in one session and returned to after a deliberate gap, rather than completed in one sitting. Generation sessions are AI-free by rule. The Playbook’s “cognitive gym” (idea 35) — deliberately unassisted practice, defended pedagogically rather than punitively — is the formal home of this element.
The agents. None, by design.
The boundary. Silence is scheduled support, not abandonment. The student is alone with the problem on purpose, for a defined period, with a mentor waiting at the end of it.
What it does. Thinking independently is one skill. Holding the result is another. As soon as someone starts thinking their own way, the environment tends to push back — first with ridicule, then suspicion, then isolation — and the most able people learn to hide exactly what is most valuable about them.
The evidence. Original work is punished before it is rewarded (Wang, Veugelers and Stephan). Original people buy low and sell high in ideas, which requires tolerating a holding period (Sternberg and Lubart). Independence of judgement and preference for complexity mark original people (Barron; Feist). Amabile’s componential theory makes intrinsic motivation central precisely because external pressure is what wears originality down. The Original Thinking Framework treats the Nerve as trained socially, through repeated and survivable exposure — and warns, via Kyaga’s registry data, that it needs scaffolding.
At CTU. Public defences in which a well-argued position that turns out wrong can still score well; explicit reward for instructive failure; cohorts large enough that nobody holds an unpopular idea alone. The studio also teaches the most important distinction a young thinker can learn: substantive criticism improves an idea; social pressure only wants it silenced. A crit that separates judgement of the work from judgement of the person teaches that distinction by practice. Resilience is not trained in isolation; it is built in a community pulling in the same direction.
The agents. The Sparring Partner produces the strongest case against a finished draft, so the student meets real opposition in a safe setting first. The Red-Team treats a predicted novelty penalty as information about timing and framing, never as a reason to stop.
The boundary. Nobody is graded on agreement with the examiner. Examiner calibration (Playbook idea 19) is the safeguard.
The studio costs contact time. The agentic university produces contact time. The strategy works only if the second is committed to the first — and if the agents inside the studio are configured to widen thinking rather than narrow it.
The Agentic University specifies eight agent archetypes. Four of them are, in effect, hour-recovery machines:
The Course Concierge absorbs logistics — deadlines, rules, the question asked forty times a semester.
The Gateway Tutor absorbs repeated explanation in the courses where students most often fall behind.
The Feedback Agent absorbs first-pass formative feedback, never summative marking.
The Instructional Design Agent drafts syllabi, slides and exercises for human authorship.
Georgia Tech’s 500-plus saved teacher hours came largely from exactly this kind of work. The Playbook’s idea 14 names the choice that follows: recovered hours become either better teaching or a staffing cut, and that choice is the difference between an agentic university and an automated one. The rule this report proposes is the Playbook’s, made specific: before the savings materialise, commit in writing that recovered hours go to studio, critique, defence and mentoring. The same passage names the trap. The supervision loop that makes the agents work — the agent steward sampling conversations each week — is the line item cut first, and nothing visibly breaks for months.
Inside the studio, a different configuration applies. The Studio Critic already exists in the architecture. To it, this report adds a student-facing version of the six agents specified in ENSI’s Originality Engine, each with the rule that keeps it from homogenising:
The Scout sends each student or cohort a weekly dispatch of mechanisms from fields they have never worked in, with a distance quota and no relevance ranking — ranking is a convergence operation.
The Sparring Partner sees work only after a complete human draft exists. It criticises and never rewrites, and it always argues the strongest opposing case rather than a balanced review.
The Divergence Auditor works for the programme, not the student’s grade. Each semester it measures how far theses and projects sit from the field’s mainstream, and how far they sit from each other, and flags cohorts where output rises while distance shrinks — the signature of creeping AI dependence. It is never visible while students are drafting, and it never becomes a mark.
The Blender takes a problem stated as a structure — entities, relations, constraints — and returns analogies from remote domains, justified relation by relation. It returns mappings, not solutions.
The Archivist maintains each student’s idea portfolio and tags every item as human-generated, AI-assisted or AI-produced. That tagging is what later makes it possible to assess own thought at all.
The Red-Team forecasts how examiners, reviewers or industry partners will receive a novel idea, and where its novelty will be misjudged.
Five rules from the Originality Engine govern the layer, with one addition for a university:
Humans generate first. No agent contributes content before a complete human first attempt exists.
AI diverges; humans converge. When agents do produce options, they are asked for the unusual ones; selection and synthesis stay with the student.
Never accept the model’s first framing. The student reframes at least once, unaided, before any agent proceeds.
Measure drift every semester. Convergence is gradual and invisible from inside; it shows only in the numbers.
Provenance or it did not happen. Every artefact carries its human/AI tag from the start.
Students always know which layer they are in. The explanation layer answers questions; the studio layer only asks them. Mixing the two is how the tutor’s helpfulness leaks into the studio.
The same operating model applies as for every other agent. A named agent steward samples studio-agent behaviour, because instruction dilution will pull the Sparring Partner toward rewriting just as it pulled CS50’s duck toward handing out code. The evidence lead evaluates the studio agents like any other deployment. And two CTU-specific cautions from the Playbook carry over: the layer should be CTU-built and grounded (idea 9), with the interaction stream retained by the university (idea 21), and frontier models serve Czech technical language measurably worse than English (idea 32), so studio agents working in Czech need their own evaluation before they are trusted.
The four outcomes in The Agentic University are defensive — they protect graduates and their clients from machine error. The fifth is generative: it certifies that the graduate can supply what the machine cannot.
The Playbook’s idea 5 proposes that every CTU programme map four outcomes to named assessments: verification, frontier judgement, unassisted core reasoning, and accountability for results one did not personally generate. Each traces to specific evidence, and together they describe an engineer who can safely direct and check machine systems. None of them describes an engineer who can find the problem worth pointing those systems at.
This report proposes adding a fifth:
Origination — the graduate can find a problem worth solving in their discipline, construct it deliberately, generate beyond the obvious answers, and hold a reasoned position on it under informed criticism. Consistent with the standard definition, it is assessed on originality and effectiveness together: the position must be new to the setting and it must work.
The four and the fifth depend on each other. Origination without unassisted core reasoning is improvisation without substance; verification without origination produces an excellent checker of other people’s ideas. A graduate with all five can do what the labour-market evidence says is becoming scarce: decide what to build, direct machines to build much of it, and know when the result is wrong.
The route is the one the Playbook already specifies for the other four, and for the same reason: accredited learning outcomes are the only teaching change that survives a change of dean, a budget round, or the departure of the enthusiast who started it. Origination belongs in the EuroTeQ Framework of Qualifications rather than in a CTU-only scheme, and it goes through the national accreditation methodology at each programme’s next reaccreditation.
Three instruments make it assessable rather than decorative:
The rejected-problem record. For every major project the student documents the problem framings they considered and why they chose one. This is the Original Thinking Framework’s behavioural test turned into a deliverable: if there was never a rejected problem, no problem was found.
The defended position. An oral defence, before examiners who did not supervise the work, of a claim the student originated.
The working artefact. Something that runs — a prototype, a model, a design that survives analysis — so that effectiveness is judged by reality rather than by impression.
The failure mode is the one the Playbook names for all graduate-attribute schemes: documentation theatre, the outcome present in the file and absent from the room. The countermeasure is the same — an outcome without a named assessment instrument is not an outcome — plus one more. The Playbook’s Removal Register (idea 30) requires every programme to name what it retired. Origination needs room in a crowded degree, and that room must come from somewhere specific, usually from content now delivered better by the explanation layer.
Start where the Playbook suggests starting the mapping exercise: three pilot programmes, one each from FIT, the Faculty of Mechanical Engineering and the Faculty of Architecture. The honest first result will be that few programmes can currently point to any assessment of origination. Better to learn that on three programmes than on 221.
Assessment decides what students actually practise. If the only verdicts reward the expected answer, no amount of studio rhetoric will produce original thinkers. The design problem is to certify originality without making it too risky to attempt.
The foundation is the Playbook’s highest-scoring idea, two-lane assessment at programme level (idea 1), grounded in TEQSA’s assessment-reform work. A small number of secured points per programme certify what must be certified under controlled conditions — above all, unassisted core reasoning. Everything else is developmental and open. This report adds a specification for the open lane: it is where origination is formed and assessed, under the studio rules of section 8.
Four instruments carry the weight:
Assessment twins (Playbook idea 7) pair an open project with a short secured oral on the same outcome, so the student must show that the thinking in the artefact is theirs.
Defence Week (idea 16) — a fixed institutional period of oral examination with cross-department examiner pools — becomes, in the upper years, a week of idea defences: students present a problem they found and a position they hold, and examiners attack it.
The AI-native capstone (idea 22), assessed on design decisions, verification and defence rather than authorship, gains one requirement: the problem is found by the student, and the rejected-problem record is part of the submission.
Process portfolios (idea 34) — assessing the trajectory of work rather than only the endpoint — suit studio and thesis work, and the Archivist’s provenance tags make them credible.
What gets graded matters as much as how. Distance from the obvious is not merit. Runco and Jaeger’s definition binds originality to effectiveness, and a maximally strange answer is usually noise. Examiners grade the quality of the framing, the reasoning behind the choices, the handling of objections and whether the result works — not how unusual it sounds.
Two design principles protect the Nerve. First, most studio work is ungraded. The Learning Exposure evidence says error affordability is what permits repetition; a studio in which every attempt counts toward a mark teaches caution, not originality. Certification happens at a few defined points. Second, a well-defended position that turns out wrong can still earn a strong grade, provided the reasoning was sound and the student can say what the failure taught them. Metcalfe’s finding that errors followed by correction produce strong learning is the pedagogical justification.
Fairness needs explicit attention, because oral and open assessment introduce biases that written examination partly hid. The detector evidence is the warning: Liang and colleagues found GPT detectors misclassifying roughly 61% of non-native English speakers’ essays as AI-written, because they read fluency as authenticity. Human examiners carry versions of the same bias — reading confidence as competence, polish as originality, and quietness as emptiness. Students working in a second language, introverted students and neurodivergent students are the most exposed. The countermeasures are specific: grade the reasoning and the defence of decisions, not the style; calibrate examiners by double-marking samples and measuring variance (idea 19); publish outcome disparities between groups; and never reintroduce AI detection informally after withdrawing it formally (idea 13).
A strategy becomes a priority when it changes what the institution pays for, whom it promotes, and who is answerable. Seven levers do that here, and most of them already exist in the Playbook.
CTU is unusually well placed to make this move. Its two most-cited research topics of the last five years are artificial intelligence and machine learning. CIIRC runs ROBOPROX, EDIH CTU and the AI-MATTERS facility. Since February 2026 the university has been led by a rector who is a professor of artificial intelligence. And its Strategic Plan 2021+ already commits it to raising the quality and success of study (goal 1.2) and bringing practice into teaching (goal 1.3). The mandate exists; what is missing is the operating decision.
The seven levers:
The hour-conversion rule. Recovered teaching hours go to studio, critique, defence and mentoring — committed in writing before the savings arrive (idea 14).
Promotion credit. Evidenced teaching redesign, explicitly including studio teaching, counts in internal evaluation and faculty promotion criteria (idea 6), with a real evidentiary bar: a redesigned course, a pre-registered evaluation, a published result.
Redesign fellowships with an origination deliverable. Twenty a year, a semester of released time each (idea 17); every fellowship delivers a course in which students find, build and defend something of their own.
A named studio lead in each faculty. The Agentic University requires four named roles — course owner, agent steward, evidence lead, certifying academic — because deployments without named owners stall. The studio needs the same: one academic per faculty accountable for studio cohorts, mentor allocation and the idea defences.
The Teaching Evidence Unit (idea 10) treats every studio pilot as a study, with a comparison condition and pre-registered outcomes, reporting to academic governance rather than to the programme it evaluates.
Structural funds. The Playbook notes that a workstream written into the AIML Research Centre and AI European Centre of Excellence proposals at drafting stage gets funded, while the same workstream added later comes out of the teaching budget (idea 33). The studio belongs in those drafts now.
Export only after delivery. The EuroTeQ alliance, with some 115,000 students, is the natural route to make origination a European engineering outcome (idea 36) — but only once CTU has results. Announcing leadership before building capability is the most reliable way to discredit the programme.
The fifth lever deserves emphasis, because studio pedagogy is exactly where evidence theatre flourishes. Students enjoy studios; satisfaction surveys will be excellent and will say nothing about whether anyone learned to think. The study CTU can run is straightforward in design: studio and conventional sections of the same course, compared on origination instruments, on secured-lane performance in the foundations, and on progression. One hypothesis deserves particular care. The Learning Exposure work notes that meaning is what keeps a person on a path long enough for anything to happen, which suggests early ownership of a real problem might reduce first-year dropout. That is plausible and untested. It should be run as a hypothesis, not announced as a result.
Originality is hard to measure and easy to fake. The answer is a small portfolio of imperfect gauges read together at programme level — never a single score, and never an individual’s grade.
Five gauges, adapted from the Originality Engine‘s measurement layer and the Learning Exposure instrument:
Origination rate — the share of graduates who defended a problem they found themselves, in a secured or public defence.
Problem-finding share — the proportion of upper-year project time spent constructing and choosing problems rather than executing pre-framed ones. The Originality Engine sets its floor at one hour in five; most degrees run close to zero.
Divergence trajectory — the semantic distance of theses and capstones from the field’s mainstream and from each other, cohort by cohort, using the semantic-distance logic of Beaty and Johnson’s SemDis work. The warning signal is rising output with shrinking distance. It is read alongside quality, never alone.
Feedback ecology by programme — reality contact, feedback latency and error affordability, scored annually with the Learning Exposure instrument, so that a programme can see whether its environment is becoming more truthful.
Studio contact — hours of small-group critique and mentoring per student, tracked against the hours the agents recovered, so the hour-conversion rule can be checked.
Three guard metrics sit beside them, to make sure the floor is not traded for the ceiling: the first-year failure rate, pass rates in the secured lane, and subgroup disparities in defence outcomes.
The limits should be stated up front. Creativity measurement has a long, contested history; the Original Thinking library documents six decades of argument over what divergent-thinking scores predict, and Simonton’s work shows how carefully indicators of scientific creativity must be handled. Each gauge above is a proxy, each can be gamed, and each degrades once it becomes a target. The defence is the portfolio: several weakly related measures, read together, at programme level, with academic judgement as the final instrument. The gauges exist to help the institution notice, not to optimise a number.
The sequence follows the Playbook’s lesson that ranking and sequencing come apart: cheap enabling moves first, accreditation changes last.
First ninety days
Publish the hour-conversion commitment, before any savings are claimed.
Choose three pilot programmes (FIT, Mechanical Engineering, Architecture) and form one cross-faculty studio cohort in each.
Agree the evaluation protocol with the Teaching Evidence Unit and record baselines for all eight gauges.
Draft the origination outcome and its three instruments.
Write the studio workstream into the structural-fund proposals currently in preparation.
Source the first real problems from one industrial partner and from existing talent environments such as FIKS and FEL Camp.
Year one
Run the pilot studios against pre-registered comparison sections.
Put the Studio Critic and Sparring Partner into service under a named agent steward, with the six layer rules enforced.
Hold the first Defence Week, including idea defences in the pilot programmes.
Run the first round of redesign fellowships, each with an origination deliverable.
Draft the promotion criterion at faculty level, where CTU has unilateral control.
Build the mentor pool, including doctoral students, and pilot the Junior Engineer Compact with one faculty, one partner and twenty students.
Years two and three
Map origination to named assessments in the pilot programmes’ accreditation files, then in every programme at its next reaccreditation.
Extend studio cohorts to all eight faculties where the pilot evidence supports it — and redesign them where it does not.
Publish the gauges and the trial results, including the failures.
Take origination to EuroTeQ as a proposed European engineering outcome, once CTU has delivery and results to show.
Each failure below is predictable from the evidence, and each needs its countermeasure built in from the start.
1. Creativity theatre. The most likely failure. The university announces an innovation week, buys beanbags, runs hackathons and counts participants. The Original Thinking Framework’s opening point is that originality treated as a mood rather than a practice does not develop. Countermeasure: the studio runs on named disciplines with repetitions, reality contact and fast feedback, and the Evidence Unit measures learning, not enthusiasm.
2. An elite track while the floor gives way. A studio for the top students, while nearly a third of first-years fail, would be both unjust and unpersuasive. Countermeasure: sequence and scope. Gateway tutoring (Playbook idea 2) comes first; the studio is for every student at their own level, as the Four C evidence allows; the first-year failure rate is a standing guard metric.
3. The studio makes students think alike. If agents enter generation sessions, the studio becomes the most efficient convergence machine in the building. Countermeasure: the six layer rules, the Divergence Auditor at programme level, and an agent steward who samples studio-agent behaviour weekly.
4. Fluency and confidence mistaken for originality. Open and oral assessment can reward the articulate and penalise the quiet, second-language and neurodivergent students who may hold the most original ideas. Countermeasure: grade reasoning and decisions rather than style, calibrate examiners, and publish disparities.
5. The institution punishes the Nerve it says it wants. This is the hardest failure because it is cultural. Rubrics penalise the unexpected, examiners reward agreement, and where standing out is quietly resented — ambition read as arrogance, enthusiasm as naivety, disagreement as insult — capable students learn to hide their best ideas. Countermeasure: reward well-defended failure, draw examiners from outside the supervising group, make the rectorate’s signal explicit and repeated, and build cohorts in which nobody stands out alone.
6. Metric capture. Once origination rate or divergence becomes a target, programmes will manufacture it. Countermeasure: programme-level reporting only, a portfolio rather than a single score, and academic judgement retained as the final instrument.
One problem remains that this strategy does not solve. The Agentic University called the hollowed apprenticeship — graduates who perform like juniors in a market with fewer junior roles — the open problem, and it is still open. Origination makes graduates better placed to create roles rather than wait for them, and the Junior Engineer Compact moves part of the apprenticeship into the degree. Neither fixes the labour market. A university that has thought about it for three years will be better placed than one that has not.
The argument of this series can now be stated end to end. Explanation has become cheap, so the agentic university can buy back the scarce hours of experienced people. Those hours are worth most where machines are weakest and degrees are thinnest: finding problems, holding positions, and thinking by making things that must work. Spent there, deliberately, under honest feedback, they produce the one graduate the next decade will pay a premium for — a person who can decide what is worth doing, direct machines to do much of it, and tell when the result is wrong.
That changes how an institution treats its most able students. At most universities they are tolerated exceptions — accommodated if they are quiet, managed if they are not. A university built around origination treats them as the people it is built around, and extends the same treatment to every student at their own level. It builds deliberately what usually exists only by accident: peers, public judgement, mentors, real problems, rewarded teachers, silence, and a community that makes standing alone less lonely.
A country’s progress is not only an economic question. It is also a question of how many people grow up believing that their own idea has value and is worth finishing. That cannot be bought or imported. It can only be built, one person at a time — and a technical university that already leads Europe in the science of machine intelligence is the natural place to show how.
Knowledge can be licensed. Explanation can be rented for $1.50 a student. A person who can find the problem nobody assigned, build an answer that works, and hold it when the room disagrees cannot be bought from any vendor. That is the last human monopoly — and producing those people is now what a university is for.