AI for Teaching Engineering: What Actually Works

September 23, 2026
blog image

Written by the ENSI Foresight Division on a library of 107 primary documents — randomised trials, meta-analyses, regulator guidance, institutional reports and deployed-system evaluations — downloaded and indexed in the “AI for Teaching at CTU” library, and cross-read against ENSI’s Education for the Agentic Age library.

Every university in Europe is currently running the same conversation, and it is the wrong one. The conversation is about permission: what may students use, what must they declare, how do we catch them. It produces policy documents, disclosure forms and detection subscriptions, and it produces almost no change in what is actually taught or how anyone actually learns. Meanwhile the evidence base — which has grown from nothing to several hundred studies in three years, and which this library assembles — has been quietly answering a different and far more consequential question: what does a machine that can explain anything, at any hour, at negligible cost, do to the production of an engineer?

The answer that emerges is not the one either camp expected. The optimists expected a tutor in every pocket and a two-sigma revolution; the pessimists expected a generation that cannot think. The measured reality is sharper and more useful than either. Where AI tutoring has been put through a proper randomised trial with a pedagogically constrained system — Kestin and colleagues’ Harvard physics experiment, published in Nature Scientific Reports, where students learned more than twice as much in less time than in an expert-taught active-learning class; the World Bank’s six-week Nigerian trial returning 0.31 standard deviations, among the most cost-effective learning interventions ever measured; Stanford’s Tutor CoPilot RCT across 900 tutors — the effects are large, real, and replicated across wildly different contexts. And where AI has simply been made available as a general-purpose assistant with no pedagogical constraint, the measured effect on thinking runs the other way: Microsoft Research and Carnegie Mellon’s study of 319 knowledge workers found that higher confidence in the AI predicts less critical-thinking effort, with users shifting from producing judgement to verifying someone else’s.

Those two findings are not in tension. They are the same finding stated twice. The variable that decides whether AI raises or lowers human capability is not the model — it is the structure wrapped around the model. Kestin’s tutor was explicitly forbidden to give answers; it was built to withhold, to question, to force retrieval. CS50’s duck at Harvard, the largest educational agent deployment on record at 211,000 students and 10 million queries, was engineered with the same intention — and its own published evaluation reports that 22% of responses contained code blocks despite instructions not to hand out solutions, rising to 48% at conversation level. That is the whole discipline in one number: the pedagogy is in the guardrail, and the guardrail leaks. An institution that buys models and skips the structure has bought the failure mode without the benefit.

This is why “should we allow it” is the wrong conversation for a technical university in particular. The prevalence question is settled: the HEPI/Kortext survey of UK students found 92% using generative AI, up from 66% a year earlier, and 88% using it in assessed work. MIT’s own institution-wide survey, reported by its Ad Hoc Committee in August 2026, found 46% of undergraduates using LLMs daily — and, more damningly, that while two-thirds of students believed AI would matter in their careers, only 25% felt MIT was adequately preparing them to use it. That gap is the actual institutional failure. It is not a discipline problem. It is a curriculum problem, an assessment problem, and above all a staff-capability problem, and every month spent litigating permission is a month not spent closing it.

There is a further reason a technical university cannot treat this as a general higher-education issue with a general higher-education answer. The disruption is not uniform across disciplines; it is concentrated almost exactly where a technical university lives. Computing education has the deepest evidence base and the sharpest dislocation — the ITiCSE working group’s landmark survey, Becker and colleagues’ “Programming Is Hard — Or At Least It Used To Be”, and the GitHub/Microsoft/MIT randomised trial in which Copilot users completed a programming task 56% faster — because code is the modality large language models are best at. The graduate attribute that a Faculty of Information Technology has spent decades certifying is precisely the one whose market price is moving fastest. Meanwhile Stanford’s Digital Economy Lab, tracking millions of payroll records, finds employment declining specifically for young workers in the occupations most exposed to AI — the entry-level software and technical roles that are the destination of a technical university’s undergraduates. The institution’s product is being repriced from both ends simultaneously.

And yet the same discipline structure that creates the exposure also supplies the defence, which is the most under-appreciated point in this entire literature. Engineering education already possesses, as its native pedagogical form, the thing every other faculty is now scrambling to invent: assessment against a physical or functional artefact that must actually work. A bridge calculation is checked by statics, not by an examiner’s impression of the prose. A circuit either oscillates or it does not. A robot either completes the task or falls over. CDIO, the international engineering-education framework to which the project-based assessment literature in this library belongs, has been building this for twenty-five years; the CESAER white paper on the engineer of the future, written by the very association of European technical universities that a school like ČVUT belongs to, argues for exactly this competence-and-challenge-based direction. Where the humanities must now reconstruct authenticity from first principles, engineering mostly has to stop drifting away from it — away from the worksheet, the boilerplate lab report, the individually-submitted problem set that a model completes in nine seconds.

So the reframe that organises this report is this. AI in teaching is not a tool question, it is a capability question — and the capability being built or destroyed is the institution’s, not the student’s. Universities that treat generative AI as software to be procured and policed will get the measured downside: offloaded thinking, unenforceable rules, an integrity arms race they lose, and graduates who use models badly because nobody taught them to use models well. Universities that treat it as a redesign of the teaching production function — what is assessed, what is taught, what staff can do, what data the institution keeps and what it is allowed to keep — get the measured upside, which is genuinely large, and get it disproportionately for their weakest students, which is where a technical university’s real losses are concentrated. The evidence for that last claim is the most decision-relevant thing in this library, and it is finding number two.

The findings in brief

  • The tutoring effect is real, large and replicated — but only for systems built to withhold. Harvard’s constrained physics tutor beat expert-led active learning by more than 2×; Nigeria returned 0.31 SD; the pooled meta-analytic effect of ChatGPT on learning is g = 0.670. Unconstrained chatbots do not reproduce it.

  • The gains concentrate in the weakest performers. Novices gained +34% in the NBER field study versus almost nothing for experts; Tutor CoPilot helped students of lower-rated tutors most. For an institution with 30%+ first-year failure, this is the single highest-value finding in the library.

  • AI-text detection does not work, and building policy on it is an equity liability. Fourteen detectors failed systematic testing; GPT detectors flagged ~61% of non-native-speaker essays as AI-written. Assessment redesign is the only durable lane.

  • Computing and programming education is the most disrupted subject on earth and has the best evidence about it — and the disruption is to learning objectives, not merely to cheating.

  • Over-reliance is measurable, and it is the real cost — not plagiarism. Confidence in AI predicts reduced critical-thinking effort; students know it and are frightened of it (90% of MIT respondents concerned about overreliance).

  • A deployed teaching agent is astonishingly cheap and demonstrably useful — and it leaks. CS50: $1.50 per student per year, 94% found it helpful, 22% of responses handed out code anyway.

  • Faculty capability, not technology, is the binding constraint — and the frameworks to fix it (UNESCO, DigCompEdu, ETH Zurich’s lecturer framework) already exist and are unused.

  • Engineering’s artefact-based pedagogy is a structural advantage nobody is exploiting — the lab, the design studio and the capstone are already AI-resistant assessment.

  • The law has already decided most of this is high-risk. The EU AI Act’s Annex III names admission, outcome evaluation, level assignment and exam monitoring as high-risk uses — which is most of what a university would want to automate.

  • Almost none of it has been evaluated to the standard the university would demand of its own research — and early-warning models specifically degrade across cohorts and subgroups.

How this report is organised

The ten findings are ranked by strength of evidence and decision-weight combined — how well the claim is established, and how much changes if you act on it. Four tests decide the ranking: whether the core result rests on randomised or quasi-experimental designs rather than surveys of opinion; whether independent teams in independent contexts converge on it; whether the outcome measured is learning or performance rather than satisfaction; and whether anything has been replicated at scale in a real institution rather than a laboratory. A result from a pre-registered RCT with a learning outcome outranks a beautifully-argued position paper. A deployment evaluation with published failure rates outranks a vendor case study. Sector surveys of what administrators intend rank last, however often they are quoted.

Each finding follows the same discipline: the claim in one line; the actual studies from the library that support it; an explicit statement of what the evidence does not show, because this field is drowning in over-claiming in both directions; and the institutional move it obliges. Where a finding bears specifically on a mid-sized European public technical university — the case that motivated this library — that is drawn out at the end of the section rather than allowed to colour the evidence. The companion report, The ČVUT Playbook, does the institution-specific work in full.

1. The tutoring effect is real and large — but only for systems built to withhold

Where AI tutoring has been tested properly, the effect sizes are among the largest in the history of education research. Every one of those trials used a system deliberately engineered not to answer.

Start with the strongest single study in the library. Kestin, Miller and colleagues ran a randomised crossover trial in Harvard’s PS2 physics course, published in Nature Scientific Reports in 2025. Students learned the same material either in a class taught by expert instructors using active-learning methods — already the gold standard against which everything else underperforms — or with an AI tutor built on GPT-4 and constrained by a detailed pedagogical prompt. Students in the AI condition learned more than twice as much, in less time, and reported higher engagement and motivation. The comparison is what makes it remarkable: this is not AI versus a bad lecture. It is AI versus the best-evidenced form of human undergraduate physics teaching, in one of the world’s most selective institutions, and the AI condition won on a pre-specified learning measure.

The result does not stand alone. The World Bank’s From Chalkboards to Chatbots working paper reports a six-week randomised trial of GPT-4-based virtual tutoring in Nigeria returning 0.31 standard deviations — which the authors, using standard learning-adjusted year conversions, place at the equivalent of one and a half to two years of business-as-usual schooling, and among the most cost-effective interventions ever measured in the education literature. Stanford’s Tutor CoPilot trial — the first RCT of a human–AI tutoring system in live tutoring, across 900 tutors and 1,800 students — found +4 percentage points of topic mastery overall at roughly $20 per tutor per year. An exploratory randomised trial in authentic UK classrooms, included here as the closest European deployment evidence, tested both learning effect and safeguarding simultaneously. And the meta-analytic anchor pools 35 experimental studies across 4,193 participants and reports g = 0.670 for ChatGPT’s effect on learning outcomes, with no significant publication bias detected.

For context on how good that is, the library carries both of the necessary baselines. Bloom’s 1984 paper — the source of the two-sigma claim that every AI-tutoring pitch deck implicitly invokes — established one-to-one human tutoring plus mastery learning as the ceiling nobody could afford. The pre-LLM reality check is the Ma, Adesope, Nesbit and Liu meta-analysis in the Journal of Educational Psychology: 107 effect sizes across 14,321 learners found intelligent tutoring systems outperforming large-group instruction and non-ITS computer instruction, but not outperforming individualised human tutoring or small-group instruction. Thirty years of intelligent tutoring systems bought a real but modest gain at high engineering cost per subject. The current generation is producing larger effects, in weeks, in domains nobody hand-authored.

Now the constraint that makes the whole finding conditional. In every trial above, the system was built to withhold. Kestin’s tutor was instructed to reveal one step at a time, to ask before telling, to refuse to produce the final answer. Tutor CoPilot does not tutor the student at all — it coaches the human tutor in real time. CS50’s duck at Harvard, whose own published evaluation is the most honest deployment document in this library, was likewise engineered to refuse solutions. There is no trial in this library in which handing students an unconstrained frontier model and letting them chat produced a learning gain of this magnitude. The pedagogy is not in the model. It is in the wrapper — the system prompt, the retrieval over the actual course material, the refusal policy, the scaffolding sequence.

What the evidence does not show. It does not show that AI tutors beat good teaching in general — Kestin’s design gave the AI condition a highly structured, expertly-prompted tutor in a single well-defined physics topic, which is a favourable case and was designed to be. It does not show durability: almost every study measures learning at or near the end of the intervention, and the library contains no strong evidence on retention months later. It does not show that gains survive when the tutor is built by a busy academic rather than a research team. And the meta-analytic g = 0.670 pools studies of highly variable quality with mostly short interventions and mostly proximal outcome measures — the standard conditions under which education effect sizes later shrink.

What it obliges. Stop procuring chatbots and start commissioning constrained tutors, subject by subject, grounded in the institution’s own course material, with the refusal behaviour specified as a pedagogical requirement rather than a safety afterthought. The unit of investment is not a licence. It is a course.

2. The gains concentrate in the weakest — and that is where a technical university’s losses are

Across every well-designed study in this library, the benefit of AI assistance is largest for the least expert and smallest for the most expert. For an institution losing a third of its first-year cohort, this is the most valuable finding here.

The pattern is extraordinarily consistent across domains that share nothing else. Brynjolfsson, Li and Raymond’s NBER field study of 5,179 customer-support agents found an average productivity gain of +14% issues resolved per hour — but +34% for novice and low-skilled workers, and minimal impact on experienced high-skilled workers. Their interpretation, which the data supports, is that the model diffuses the tacit knowledge of the best performers to everyone else. Stanford’s Tutor CoPilot trial found the same shape inside education: +4pp overall, but +9pp for students working with lower-rated tutors — the AI closed part of the gap between a weak tutor and a strong one. The Dell’Acqua, Mollick and Lakhani field experiment at BCG, run on 758 consultants, found the largest gains among below-average performers, who improved dramatically toward the group mean.

This is a compression effect, and it points somewhere very specific. A technical university’s most expensive, most persistent and least discussed failure is not the mediocrity of its best students. It is attrition in the first two years, concentrated in the gateway subjects — mathematical analysis, physics, mechanics, the first programming course — where a student who falls two weeks behind cannot recover because the only remediation available is a queue outside an office hour that clashes with another lecture. The published Czech numbers make the scale concrete: at ČVUT, first-year bachelor study failure ran at 31.8% university-wide in 2024, ranging from 9.5% at the Faculty of Architecture to 51.1% at the Faculty of Mechanical Engineering and 48.1% at Transportation Sciences. Half of one faculty’s incoming bachelor cohort fails in year one. Those are not students who lack capacity; they are overwhelmingly students who lacked a patient explanation at eleven at night in week six.

That is precisely the good that a constrained tutor supplies at negligible marginal cost, and it is precisely the population the evidence says benefits most. It is worth being blunt about the arithmetic, because it reframes the entire investment case. A technical university spends heavily on admissions marketing to fill a cohort, then loses a third of it to a failure mode that a well-built tutoring layer measurably addresses. The retention gain is worth more, in both money and mission, than every productivity saving in the administrative use cases that dominate university AI strategies.

Two cautions before anyone builds the business case. First, compression is not only good news: if AI lifts the floor without lifting the ceiling, the signal value of a degree compresses too, and the strongest students gain least from the institution’s investment. Second, the same compression logic that helps a struggling first-year is what makes over-reliance dangerous later — a student permanently held at the level the tool provides has been given a floor and a ceiling in the same object. Finding five deals with that directly.

What the evidence does not show. None of the compression studies are from engineering education specifically, and the two largest are from work settings rather than degree programmes. They measure task performance, not the acquisition of durable expertise — and the mechanism, diffusion of expert tacit knowledge to novices, is precisely the mechanism that could substitute for learning rather than produce it. No study in this library demonstrates that AI tutoring reduces university dropout. That trial has not been run, which is itself a finding, and it is exactly the trial a technical university with a 31.8% first-year failure rate is best placed in Europe to run.

What it obliges. Point the first serious deployment at the gateway courses with the worst failure rates, not at the flagship master’s programmes where the enthusiasts are. And instrument it as a trial, so the institution ends up owning evidence rather than anecdote.

3. Detection does not work — assessment redesign is the only durable lane

AI-text detection has been systematically tested and it fails; worse, it fails asymmetrically against non-native speakers. Any integrity policy resting on detection is both ineffective and an equity liability.

Weber-Wulff and colleagues, working through the European Network for Academic Integrity, tested fourteen AI-text detection tools under controlled conditions. The finding is unambiguous: the tools are neither accurate nor reliable, they skew toward classifying output as human-written, and their performance collapses under light obfuscation — machine translation, minor paraphrase, a pass through another model. This is not a maturity problem that a better product cycle fixes; it is close to information-theoretically inevitable as models converge on fluent, unremarkable prose.

The equity finding is worse and should end the argument on its own. Liang, Yuksekgonul, Mao, Wu and Zou at Stanford ran GPT detectors over TOEFL essays written by non-native English speakers and over essays by native-speaking US eighth-graders. The detectors were near-perfect on the native-speaker writing and misclassified roughly 61% of the non-native-speaker essays as AI-generated, with over half flagged by all seven detectors tested. The mechanism is that detectors key on lexical richness and syntactic variety — exactly the dimensions on which a competent second-language writer differs from a native one. For any technical university with a substantial international cohort, deploying such a tool means systematically accusing international students of misconduct at several times the rate of domestic ones, on the basis of their second-language fluency.

So the enforcement lane is closed. The library’s answer to what replaces it is unusually well-developed, because Australia’s regulator did the work first. TEQSA’s 2023 discussion paper Assessment Reform for the Age of Artificial Intelligence set out the two principles that now underpin most serious sector guidance worldwide, and its 2025 follow-up, Enacting Assessment Reform in a Time of AI, reports what institutions actually did with them. The core move is to stop treating every assessment as if it served one purpose. Some assessment exists to certify that a named human has a capability — and that requires secured conditions and identity assurance, at programme level rather than in every task. Everything else exists to develop capability — and there AI use should be open, taught and part of the point. Trying to make every assignment do both jobs is what produced the unwinnable arms race. QAA’s guidance adds the governance procedure: a four-step triage of which assessments actually need redesign, run through existing internal quality assurance rather than as a parallel emergency process.

For engineering, the practical translation is more favourable than for most disciplines, and the library supplies the specifics. The CDIO paper on project-based assessment in the era of generative AI reworks the PBL evaluation grid with explicit AI-use criteria — you grade the process, the design decisions and the defence, not only the artefact. The Integrevise research report on oral assessment gives the operating model, cost and staffing constraints of running vivas at cohort scale, which is the obvious authentication lane for a technical university and the one most often dismissed as impossible without checking the numbers. And ČVUT’s own Methodological Instruction 5/2023 already contains an activity-by-activity permitted / partly-permitted / forbidden schema — a more concrete instrument than most European universities possess, though written before the assessment-redesign literature matured and now due a revision that moves it from a rules document to a design document.

What the evidence does not show. It does not show that detection tools are useless in every configuration — they retain some signal on unedited long-form output, and Turnitin-class vendors dispute the specific error rates. It does not show that secured in-person assessment is unproblematic: TEQSA is explicit that identity assurance at programme level is hard, expensive and easy to implement badly. And there is no strong evidence yet on whether two-lane assessment actually preserves standards, because it is too new to have graduated a cohort.

What it obliges. Retire detection as a basis for misconduct proceedings, immediately and explicitly, and say why in public so that staff stop relying on it informally. Then run the triage: for every programme, identify the small number of points where certification genuinely requires secured conditions, secure those properly, and free everything else to be taught with AI in the open.

4. Computing education is the most disrupted subject on earth — and the disruption is to objectives, not to cheating

The subject a technical university teaches most confidently is the one large language models are best at. The published response from the computing-education research community is not about misconduct; it is about which learning objectives are still worth certifying.

No other discipline has responded to generative AI with this much empirical work this fast, which makes computing education the closest thing the sector has to a natural experiment. The ITiCSE working group report — twenty-odd authors, the landmark community survey in this library — covers code generation, code explanation, autograders, automated feedback, integrity and curriculum response in a single document, and its conclusion is that the pedagogical questions dwarf the disciplinary ones. Becker, Denny, Finnie-Ansley, Luxton-Reilly, Prather and Santos put the point in the title of their SIGCSE paper: Programming Is Hard — Or At Least It Used To Be. Their argument is that a large fraction of CS1’s traditional learning objectives were proxies. We never actually wanted students to be able to write a for-loop from memory; we wanted them to be able to decompose a problem, and writing the loop was how we checked. The proxy has broken. The underlying objective has not.

The industry-side evidence explains why the objectives must move rather than be defended. Peng, Kalliamvakou, Cihon and Demirer’s randomised controlled trial — GitHub, Microsoft and MIT — found developers using Copilot completed a standard programming task about 56% faster than the control group. Whatever a university thinks about AI in coursework, its graduates enter a profession where this is the baseline expectation. Certifying an ability to produce code unaided, slowly, is certifying a skill the employer will not buy.

The most useful study in this angle is the most uncomfortable. Prather and colleagues observed CS1 students actually using Copilot on a real assignment — the paper is titled, from a student quote, “It’s Weird That it Knows What I Want”. What they document is a set of genuinely new interaction pathologies: drift, where the student’s mental model of the problem quietly diverges from the code accumulating on screen; over-trust, where plausible output is accepted without verification; and a collapse of the metacognitive loop that novice programming is supposed to build. Ma, Chen and Konomi’s study of dialogue logs from a beginner Python course adds the typology, clustering student–ChatGPT interaction into four distinct usage patterns and linking each to performance — which means usage pattern, not usage volume, is the variable that matters and the thing worth teaching.

The instructional-response literature is thinner but concrete. The ASEE study on ChatGPT in programming courses documents the practice of requiring students to submit their own code alongside the AI’s and account for the difference — an assessment pattern that converts the tool into the object of study. The JITE paper supplies the instructor-side view of benefits and adverse impacts, which is what faculty development has to start from.

What the evidence does not show. There is still no strong longitudinal evidence on what happens to programming expertise across a whole degree under heavy AI use — the studies are single-course, single-semester, and mostly measure task outcomes rather than developed capability. The Copilot productivity RCT measured a well-specified task, not the messy comprehension-and-maintenance work that dominates real engineering. And the four-pattern typology is from one course at one university.

What it obliges. Rewrite CS1 and CS2 learning outcomes explicitly around decomposition, specification, verification, debugging and reading unfamiliar code — the objectives that survive — and assess those directly rather than through the broken proxy. Then treat every other engineering discipline as being roughly two years behind computing on the same curve, and start the same work now rather than waiting for its own crisis.

5. Over-reliance is measurable, and it is the real cost — not plagiarism

The strongest argument against casual AI adoption is not integrity. It is that confident use of a capable model measurably reduces the critical-thinking effort of the person using it — and students are more worried about this than their teachers are.

The central study is Lee and colleagues at Microsoft Research and Carnegie Mellon, published at CHI 2025: 319 knowledge workers supplied 936 first-hand examples of generative AI use at work. Two findings matter. Higher confidence in the AI predicts less critical-thinking effort; higher self-confidence in one’s own expertise predicts more. And the nature of the effort shifts — from information gathering and problem-solving toward information verification and response integration. That is not automatically bad: verification is real cognitive work and, done well, is exactly the skill finding ten of this report argues should be taught. It is bad when it is not done, and the study’s confidence finding says that the better the tool gets, the less likely the user is to do it.

This connects to a body of work that ENSI’s Education for the Agentic Age library treats at length under cognitive debt and cognitive offloading — the well-established finding that capability which is habitually externalised is not merely unused but progressively unavailable. In a professional formation context, that is the whole risk. An engineer who cannot check the model is not an engineer who is slower; they are an engineer who cannot tell when the answer is wrong, in a profession where being unable to tell is how people get hurt.