The ČVUT Playbook: What Czech Technical University Should Actually Do

September 10, 2026
blog image

Written by the ENSI Foresight Division on a library of 107 primary documents, including CTU’s Strategic Plan 2021+, its 2024 Annual Report, “ČVUT v číslech 2024”, Methodological Instruction 5/2023 on the use of artificial intelligence, and the international evidence base assembled in the “AI for Teaching at CTU” library.

There is a specific and uncomfortable fact about Czech Technical University that ought to organise everything it does next in this area. CTU’s two most-cited research topics of the last five years are “artificial intelligence” and “machine learning”. On five-year comparison it sits among the top five European research groups in computer vision and ninth in robotics. CIIRC coordinates ROBOPROX — 468 million CZK of Excellent Research funding — runs EDIH CTU and the AI-MATTERS Testing and Experimentation Facility, is a partner in two of the five European AI and robotics networks of excellence, and has an institute named after its founder built and opened in Jaipur. Since February 2026 the university has been led by a rector who is a professor of artificial intelligence and the founder of the AI Center at the Faculty of Electrical Engineering.

And in 2024, 31.8% of CTU’s first-year bachelor students failed — 51.1% at the Faculty of Mechanical Engineering, 48.1% at Transportation Sciences, 45.6% at Nuclear Sciences and Physical Engineering, 34.0% at Information Technology. The university’s institution-wide rulebook for artificial intelligence in teaching is a methodological instruction issued in September 2023, eight pages long, structured as a table of permitted, partly permitted and forbidden activities, written before the assessment-reform literature matured and before any of the randomised trials that now define the field had been published. Its named AI-literacy provision for students is a licensed online course produced by the University of Helsinki.

That gap — between what CTU knows about AI and what CTU does with AI in its own lecture theatres — is the whole strategic situation. It is not a criticism of the instruction, which was an early and sensible response and remains more concrete than most European universities managed. It is an observation that CTU is currently a world-class producer of artificial intelligence and an ordinary consumer of it, and that the second fact is now the more consequential one. Every technical university in Europe can license the same models. Almost none of them have a rector who built an AI research centre, a CIIRC-scale institute, a top-five European computer-vision group, and a national mandate under the Czech AI Strategy to 2030 sitting in the same building as a 51% first-year failure rate. The asymmetry is the opportunity, and it has a short half-life: the differentiation is available for about three years, after which everyone will have done the obvious things.

But attrition is only the first of two problems, and the second changes what the end of a degree is for.

Stanford’s Digital Economy Lab, tracking millions of payroll records, finds employment declining specifically among young workers in the occupations most exposed to AI — entry-level software and technical roles, which is to say the destination of a technical university’s bachelor and master graduates. Set that against the compression findings that make the tutoring case so strong: +34% for novices against near-zero for experts in the NBER field study of 5,179 workers; the largest gains for below-average performers in the BCG field experiment; +9 percentage points for students of the weakest tutors in Stanford’s Tutor CoPilot trial.

Put those together and the shape of the problem is unusual. The technology compresses the performance gap between novice and expert while eroding the jobs in which that gap was historically closed. A CTU graduate can now perform like a competent junior on day one, and may find no junior position in which to become a senior. The apprenticeship function — the first three years in industry where an engineer acquires judgement by being wrong under supervision — is migrating upstream, and the only institution positioned to absorb it is the university.

This is not a distant concern. It bears directly on what a capstone, a diploma thesis, an industrial placement and a laboratory course are for. CESAER’s Engineer of the Future white paper and the CDIO tradition already describe the direction — challenge-based, competence-based, real consequences — without naming this as the reason. Several of the highest-scoring ideas below are that argument made operational.

So the reframe this playbook argues for is not “CTU should adopt AI in teaching”. Everyone will. It is that CTU should treat teaching as an application domain of its own research, and its own students as the population on which the European evidence base gets built — against two problems at once: the third of the cohort lost in year one, and the professional formation that used to happen after graduation and no longer will. That is a claim about identity, not about procurement. A university that publishes on machine learning and runs its education on intuition and committee memory is holding two incompatible epistemologies.

Three conditions make the timing favourable, and all three are transient. The mandate already exists: Strategic Plan CTU 2021+ has four pillars, of which Pillar 1 is Study, and its goals include raising the quality and success rate of study and bringing practice into teaching; Pillar 3 commits CTU to “digitalisation of activities and operations, decision-making on the basis of data”. The evidence has arrived and points at CTU’s worst number, since every well-designed study says the benefit concentrates in the weakest performers. And the university has already run the experiment without noticing — the 2024 annual report records that the Faculty of Mechanical Engineering uses artificial intelligence to predict “at risk” status from weekly examination results and offers targeted help with exam scheduling, credited with a significant fall in failure. It is not in the strategic plan, has no institutional owner, and has never been evaluated to any standard CTU would accept from a doctoral student. The programme does not start from zero; it starts from an unowned success nobody has scaled.

How the ideas are scored

Forty interventions follow. Three dimensions, each out of 10, summed to a composite out of 30. A fourth consideration — feasibility — is deliberately kept out of the score and reported separately, because feasibility should determine sequence, not merit. Scoring hard things down because they are hard is how institutions end up with a portfolio of easy things that changed nothing.

Evidence quality (E) — how well the international library actually supports this. A 9 or 10 means multiple randomised or quasi-experimental studies converging, in comparable settings, on a learning or progression outcome. A 6 or 7 means consistent professional consensus, regulator guidance, or strong observational evidence from named deployments. A 3 or 4 means it is a reasonable inference from adjacent evidence but nobody has tested this thing. A 1 or 2 means it is a bet.

Depth (D) — how structural the change is. Does it alter the teaching production function, what a degree certifies, or who is accountable for what — or does it add a capability on top of an unchanged system? A 9 or 10 changes what CTU is for the student. A 5 or 6 changes how something is delivered. A 2 or 3 is an improvement that leaves every underlying structure intact. High depth is not automatically good; it measures leverage and risk simultaneously.

Expected outcome (O) — magnitude times probability, at CTU specifically. Not “is this a good idea in general” but “what is the realistic expected value of doing this at an institution with 18,168 students, 2,247 academic staff, 221 study programmes across eight faculties and six institutes, a 31.8% first-year failure rate, 19.7% international enrolment, and a rector who is an AI professor”. An idea with a large effect that will probably not survive institutional contact scores lower than a modest effect that certainly will.

Feasibility, reported separately as one of four labels: this year · this cycle (2–3 years) · accreditation-bound (lands at the next programme reaccreditation) · hard (requires external agreement, money that does not exist, or a cultural change with no current sponsor).

The composite is a ranking instrument, not an oracle. Two honest limitations. It rewards ideas whose effects are measurable, which biases against slow cultural moves that matter and cannot be counted. And the E scores are drawn from a library assembled for this question — a different library would move them. Both are reasons to read the reasoning rather than the number.

The scoreboard — forty ideas, ranked

Read as: E evidence quality · D structural depth · O expected outcome at CTU · = composite /30 · feasibility label.

Tier 1 — the seven that clear 24

1. Two-lane assessment at programme level — E 9 · D 9 · O 8 · = 26 · accreditation-bound. Replace MP 5/2023’s activity-by-activity permission schema with a small number of properly secured certification points per programme and everything else open and taught. Scores highest because the evidence is unusually complete on both halves — detection demonstrably fails, and TEQSA’s two-lane model has a five-year institutional track record — and because it changes what a CTU degree asserts, which nothing else on this list does as directly.

2. Gateway Tutor on the five worst first-year courses — E 9 · D 7 · O 9 · = 25 · this cycle. Constrained, course-grounded tutors on the gateway subjects at FS, FD and FJFI, run as randomised trials. Highest expected outcome on the list: four independent RCTs support the intervention, every compression study says the effect concentrates in exactly the population CTU is losing, and the target number is published annually.

3. Discipline-specific verification standards — E 8 · D 9 · O 8 · = 25 · this cycle. Define, per discipline, what counts as checking a machine-produced result — in circuits, in structures, in code, in control. Scores here because the jagged-frontier evidence shows people cannot locate the capability boundary untaught, and because this is the one graduate competence that is both load-bearing and currently taught nowhere.

4. The Junior Engineer Compact — E 7 · D 10 · O 8 · = 25 · hard. With Czech industry: a structured final-year track of supervised responsibility for real work, absorbing the apprenticeship function that entry-level roles no longer perform. The deepest idea on the list and the only one addressing the collapse of the junior pipeline. Marked hard because it requires industry agreement CTU does not yet have.

5. Four graduate AI outcomes in every accredited programme — E 8 · D 9 · O 8 · = 25 · accreditation-bound. Verification, frontier judgement, unassisted core reasoning, accountability for machine-produced results — written into programme documentation with named assessment points. Scores high on depth because accredited outcomes are the only teaching change that survives a change of dean.

6. Teaching redesign counts for promotion and habilitation — E 6 · D 10 · O 8 · = 24 · hard. The single highest-leverage move on the list and the one with the weakest direct evidence, which is why it sits sixth rather than first. Every study of stalled adoption identifies incentives; none of them tested changing incentives. Depth 10 because it is the only idea here that changes what the institution rewards.

7. Assessment twins — E 8 · D 8 · O 8 · = 24 · this cycle. Pair an open, AI-permitted component with a short secured oral or practical assessing the same outcome, scheduled close together for cross-verification. Scores just below two-lane assessment because it is the implementable pattern inside that model rather than the model itself.

Tier 2 — strong (20–23)

8. CS1/CS2 outcome rewrite at FIT and FEL — E 9 · D 7 · O 7 · = 23 · accreditation-bound. Rebuild introductory programming outcomes around decomposition, specification, verification, debugging and reading unfamiliar code. Best-evidenced curriculum change available; capped on outcome only because it touches two faculties.

9. A CTU-built, course-grounded tutoring layer — E 8 · D 8 · O 7 · = 23 · this cycle. Own the pedagogically-constrained layer rather than renting a general assistant. Georgia Tech’s 76.7%-versus-31.3% accuracy gap is the evidence that grounding, not model choice, is what works.

10. The Teaching Evidence Unit — E 9 · D 7 · O 7 · = 23 · this year. Three to four people who make every deployment a pre-registered trial. Highest evidence score on the list; outcome capped because its effect is entirely indirect — it changes the quality of every other decision rather than any student’s result.

11. ETH Zurich lecturer framework, workload-credited, by discipline — E 8 · D 7 · O 7 · = 22 · this year. Adopt rather than draft; deliver in faculty cohorts with hours credited, not added. Scores well and is capped by the honest observation that no study in the library evaluates a faculty AI-development programme against a teaching outcome.

12. The Confusion Map — E 6 · D 8 · O 8 · = 22 · this cycle. Publish, per course, where students actually get stuck, derived from the tutor interaction stream. A university has never before been able to see confusion at scale in the student’s own words; ten million CS50 queries is a map of what is hard about introductory computing that no pedagogical intuition could produce.

13. Retire detection from misconduct procedure — publicly, with the reason — E 9 · D 5 · O 7 · = 21 · this year. Fourteen detectors failed systematic testing; GPT detectors misflag roughly 61% of non-native English writers. With 3,577 international students and 78 English-taught programmes, this is a live equity exposure. Cheapest high-scoring move available.

14. Convert saved lecture hours into studio and seminar contact — E 6 · D 8 · O 7 · = 21 · this cycle. The productive use of every efficiency elsewhere on this list. Scores on depth because it is the difference between an agentic university and an automated one, and it will be decided in budget meetings rather than strategy documents.

15. Universal AI-delivered pre-matriculation bridge — E 7 · D 6 · O 8 · = 21 · this year. Scale what FIT, FEL and FJFI already run in fragments — FIKS, FEL Camp, the Preparatory Week, the Mathematical and Physics Minimum — into a universal, AI-delivered bridge for every admitted student. High outcome because it attacks the failure before enrolment, when it is cheapest.

16. University-wide Defence Week — E 6 · D 8 · O 7 · = 21 · this cycle. A scheduled institutional rhythm of oral authentication rather than per-course vivas invented by exhausted individual lecturers. Turns the staffing problem of secured assessment from a distributed impossibility into a timetabling exercise.

17. Redesign fellowships — E 7 · D 7 · O 7 · = 21 · this cycle. Pay academics released time to rebuild a specific course, with a deliverable. The Ithaka evidence is unambiguous that lack of time, not lack of willingness, is the binding constraint.

18. One governance register with a risk-class sequencing rule — E 8 · D 6 · O 7 · = 21 · this year. AI Act Annex III classification, ESG route, data flows and named owner for every system, plus a published rule that no high-risk deployment precedes an evaluated low-risk one. Protects the programme from its own enthusiasts.

19. Examiner calibration — E 6 · D 7 · O 7 · = 20 · this cycle. Use AI to measure and reduce inter-examiner variance. A larger and more measurable fairness problem than cheating, entirely unaddressed, and newly tractable.

20. Students build the agents — E 4 · D 9 · O 7 · = 20 · this cycle. FIT and FEL students build the tutoring agents for FS and FSv gateway courses as assessed coursework. Collapses the cost, produces an authentic capstone with a real user, and makes the university’s own teaching the object of student engineering. Evidence score is low because nobody has published this; depth is high because it changes who does the work.

21. Retain the interaction stream institutionally — E 6 · D 8 · O 6 · = 20 · this year. MP 5/2023 already states the fact correctly: no AI tool used at CTU is operated by CTU. Today the richest teaching-improvement dataset the university could own accrues to a vendor.

22. The AI-native capstone — E 5 · D 8 · O 7 · = 20 · accreditation-bound. Every capstone ships and publicly defends an artefact built with agentic tooling, assessed on design decisions and verification rather than authorship.

Tier 3 — worth doing (16–19)

23. Diagnostic competence map at entry — E 5 · D 7 · O 7 · = 19 · this cycle. Replace a single admission score with a per-topic gap map that routes the student to specific remediation.

24. “Machines and Judgement” spine course across all eight faculties — E 6 · D 7 · O 6 · = 19 · accreditation-bound. One shared, discipline-adapted course carrying the AI-literacy and verification content, replacing the licensed general online course currently doing that job.

25. The Course Concierge — E 8 · D 4 · O 7 · = 19 · this year. Syllabus and logistics agent. Low depth by design; it is the cheapest archetype, the fastest visible staff relief, and the safest place to learn to run any of this.

26. Teaching-AI as doctoral topics at CIIRC and the AI Center — E 5 · D 7 · O 7 · = 19 · this cycle. Solves staffing, cost and publication simultaneously by making the teaching engine a research programme rather than unpaid service.

27. The Week-Six Trigger — E 6 · D 5 · O 7 · = 18 · this year. Detect and remediate at the specific point in a cumulative course where recoverable falling-behind becomes unrecoverable.

28. Audit and scale the FS at-risk model — E 7 · D 4 · O 7 · = 18 · this year. It already runs, it is credited with a fall in failure, it has no institutional owner, and it has never been audited for cohort drift or subgroup fairness — which the dropout-prediction literature says is where these models fail.

29. Second-chance architecture — E 4 · D 7 · O 7 · = 18 · this cycle. A structured re-entry path with AI-supported catch-up for the near-miss share of the 25.71%, instead of treating failure as terminal.

30. The Removal Register — E 4 · D 8 · O 6 · = 18 · accreditation-bound. Require every programme to name what it retired this cycle. Curricula only ever accrete; adding AI content without a removal instrument produces an unteachable degree.

31. Embedded CIIRC engineer per faculty — E 4 · D 7 · O 7 · = 18 · this cycle. A semester-long residency that transfers capability rather than delivering a system and leaving.

32. Czech technical-language evaluation set — E 5 · D 7 · O 6 · = 18 · this cycle. Frontier models serve Czech technical instruction measurably worse than English. CTU has the NLP capability to build the benchmark, and no one else in the country will.

33. Write teaching-AI into the structural-fund proposals — E 4 · D 6 · O 8 · = 18 · this year. The AIML Research Centre and AI European Centre of Excellence are in preparation now. A workstream written in at drafting stage is funded; one added later is not. Pure timing value, and the window closes.

34. Process portfolio assessment — E 5 · D 7 · O 5 · = 17 · accreditation-bound. Assess the trajectory of work rather than the final artefact.

35. The cognitive gym — E 6 · D 6 · O 5 · = 17 · this cycle. Deliberately unassisted practice spaces, defended pedagogically rather than punitively — the capability on which verification skill is parasitic.

36. EuroTeQ workstream and EDIH industry microcredentials — E 5 · D 6 · O 6 · = 17 · this cycle. The export move. Worthless before delivery exists, valuable immediately after.

37. Validate the admission test against post-AI outcomes — E 5 · D 6 · O 5 · = 16 · this cycle. If AI compresses performance differences, the instrument that used to predict who succeeds may no longer predict it. Nobody has checked.

Tier 4 — scored, and not recommended now (≤15)

38. Assessment variant generation at scale — E 5 · D 4 · O 6 · = 15 · this year. Twenty equivalent exam versions, generated and human-checked. Genuinely useful — but it is a component of the secured lane, not an initiative, and promoting it to a programme invites building the tool before deciding the assessment model it serves.

39. Resit and exam-scheduling optimisation — E 5 · D 3 · O 5 · = 13 · this year. Real efficiency, no structural change, and it risks becoming the visible “AI project” precisely because it is easy and uncontroversial. Do it as operations, not as strategy.

40. Peer observation of AI-mediated teaching — E 5 · D 4 · O 4 · = 13 · this year. Sound practice, but with 2,247 academic staff it consumes exactly the senior attention the redesign fellowships need, and the evidence that observation changes teaching behaviour is weak.

Deliberately not on this list, and why. Automated summative grading, AI admissions triage and remote proctoring were considered and excluded rather than scored, because all three sit inside AI Act Annex III’s high-risk categories while CTU has no compliance track record, no evaluation function and no institutional trust built. They are not bad ideas permanently; they are bad ideas first, and the cost of that mistake is not recoverable.

Tier 1 in full

1. Two-lane assessment at programme level — 26/30

In short. Retire the activity-by-activity permission schema of Methodological Instruction 5/2023 and replace it with a programme-level architecture: a small, deliberately-chosen set of secured certification points where CTU asserts that a named human holds a capability, and everything else open, AI-permitted and taught.

The mechanism, and why it is not a rules change. The current instrument asks, of each activity, may a student use AI for this? That question has no enforceable answer, and the enforcement evidence is conclusive: Weber-Wulff and colleagues, working through the European Network for Academic Integrity, found fourteen detection tools neither accurate nor reliable and defeated by light paraphrase; Liang and colleagues at Stanford found GPT detectors misclassifying roughly 61% of essays by non-native English writers as AI-generated while performing near-perfectly on native speakers. With 3,577 international students — 19.7% of the body — from over 100 nationalities and 78 of 221 programmes taught in English, a detection-founded regime at CTU accuses its international cohort at several times the domestic rate on the basis of second-language fluency.

The two-lane model asks a different and answerable question: what is this assessment for? TEQSA’s 2023 discussion paper established the principles and its 2025 follow-up reports what institutions actually built from them. Assessment of learning certifies — and therefore requires secured conditions and identity assurance, at programme level rather than in every task. Assessment for learning develops — and there AI use is open, expected, and frequently the subject of the assessment. The unwinnable arms race came from demanding both jobs from every assignment.

Why it scores 26. Evidence 9: both halves are unusually well established — the failure of detection empirically, the two-lane model through five years of regulator-guided institutional practice. Depth 9: it changes what a CTU degree asserts to a stranger, which is the institution’s actual product. Outcome 8: high confidence of adoption because it reduces staff burden rather than adding to it, capped only by accreditation timing.

What CTU already has. More than most European universities. MP 5/2023 (ČVUT_MP_2023_05_V01, eight pages, effective 25 September 2023, issued by the Vice-Rector for Bachelor and Master Studies) already did the hard analytical work of thinking activity-by-activity about where AI use is pedagogically load-bearing — that analysis is reusable, it is the form that must change. Two of its provisions should be preserved verbatim: the warning that no AI tool used at CTU is operated by CTU and that all user–tool communication is visible to the operator, and the explicit treatment of deepfake identity modification in online examinations as a disciplinary offence. The Study and Examination Code, consolidated and effective from 1 December 2025, is the harder vehicle in which secured-assessment requirements must ultimately live.

The first move. Run the QAA four-step triage across every programme, through existing internal quality assurance rather than as an emergency parallel process — which is also what keeps it accreditable with NAÚ under the recommended procedures for preparing study programmes. For each programme the output is a single page: which outcomes require certification, at which points, under what identity assurance. Expect the honest answer to be three to five points across a bachelor’s degree, not thirty.

The failure mode. Two, both common. The first is that “secured” is implemented as surveillance — proctoring software, which is both an Annex III high-risk use and, on the evidence in this library on proctoring and disability, an accessibility liability. The second is that the open lane is declared and then quietly policed anyway, with informal suspicion migrating into marking. The countermeasure to both is measurement: publish the misconduct-rate disparity between international and domestic students annually.

Cost and owner. Vice-Rector for Studies, through faculty study committees and the Academic Senate. Drafting is cheap; the real cost is contact hours for secured oral components, which idea 16 (Defence Week) exists to make affordable. The Integrevise research report gives the staffing and cost model for oral assessment at cohort scale — compute it for CTU’s actual cohorts rather than dismissing the option on intuition.

The number that proves it wrong. If, two years in, no programme has reduced its number of assessed tasks and the misconduct disparity has not narrowed, this was a document change and not a reform.

2. Gateway Tutor on the five worst first-year courses — 25/30

In short. Constrained, course-grounded AI tutors on the gateway subjects that fail the most students — mathematical analysis, physics, mechanics, first programming — beginning at the Faculty of Mechanical Engineering, deployed as randomised trials.

The mechanism. First-year failure at a technical university has a stereotyped shape: the material is strictly cumulative, a student falls two weeks behind, week seven becomes unintelligible without week five, the only remediation is a consultation hour that clashes with another lecture, attendance stops, formal failure follows in February. Nothing in that sequence requires a human at the moment of intervention. It requires a patient, correct, course-specific explanation at eleven at night, which is the one thing this technology unambiguously supplies — provided it is built to withhold.

Why it scores 25. Evidence 9: four independent randomised trials converge — Kestin and colleagues’ Harvard physics crossover trial, where students learned more than twice as much in less time than in expert-led active learning; the World Bank’s six-week Nigerian RCT at 0.31 standard deviations; Stanford’s Tutor CoPilot at +4 points overall and +9 points for students of the weakest tutors; and a pooled meta-analytic effect of g = 0.670 across 35 studies and 4,193 participants. Outcome 9, the highest on the list, because the compression findings say the effect lands exactly where CTU’s losses are and the target metric is already published. Depth only 7 — it improves delivery of an unchanged curriculum, which is precisely why it is safe to do first.

What CTU already has. The Faculty of Mechanical Engineering already uses artificial intelligence to predict “at risk” status from weekly examination-period results and offers targeted help with exam scheduling — recorded in the 2024 annual report and credited with a significant fall in failure. CTU’s teaching-AI programme starts from an unowned success nobody has scaled. It also has a substantial existing scaffolding culture to attach to (FIKS, FEL Camp, the FJFI Preparatory Week and its free senior-student tutor system, FD’s mathematics and physics tutoring, the CIPS, ELSA and KC counselling centres), KOS and Moodle as substrate, and in CIIRC and the FEL AI Center the capability to build this properly.

The first move. Two courses, not twenty. Assemble each course corpus — lecture notes, problem sets, worked solutions, past examinations, rights-cleared textbook material — and build retrieval-grounded tutors over them with the pedagogical constraint specified as an engineering requirement: one step at a time, question before answer, never the final result to an assessed problem. Georgia Tech’s numbers are the argument for grounding over model choice: 76.7% answer accuracy against 31.3% for a generic assistant baseline on the same evaluation, and coverage climbing from 21% at 80% precision to over 96% at over 86% precision through iteration.