Written by the ENSI Foresight Division on a library of 107 primary documents, including CTU’s Strategic Plan 2021+, its 2024 Annual Report, “ČVUT v číslech 2024”, Methodological Instruction 5/2023 on the use of artificial intelligence, and the international evidence base assembled in the “AI for Teaching at CTU” library.
There is a specific and uncomfortable fact about Czech Technical University that ought to organise everything it does next in this area. CTU’s two most-cited research topics of the last five years are “artificial intelligence” and “machine learning”. On five-year comparison it sits among the top five European research groups in computer vision and ninth in robotics. CIIRC coordinates ROBOPROX — 468 million CZK of Excellent Research funding — runs EDIH CTU and the AI-MATTERS Testing and Experimentation Facility, is a partner in two of the five European AI and robotics networks of excellence, and has an institute named after its founder built and opened in Jaipur. Since February 2026 the university has been led by a rector who is a professor of artificial intelligence and the founder of the AI Center at the Faculty of Electrical Engineering.
And in 2024, 31.8% of CTU’s first-year bachelor students failed — 51.1% at the Faculty of Mechanical Engineering, 48.1% at Transportation Sciences, 45.6% at Nuclear Sciences and Physical Engineering, 34.0% at Information Technology. The university’s institution-wide rulebook for artificial intelligence in teaching is a methodological instruction issued in September 2023, eight pages long, structured as a table of permitted, partly permitted and forbidden activities, written before the assessment-reform literature matured and before any of the randomised trials that now define the field had been published. Its named AI-literacy provision for students is a licensed online course produced by the University of Helsinki.
That gap — between what CTU knows about AI and what CTU does with AI in its own lecture theatres — is the whole strategic situation. It is not a criticism of the instruction, which was an early and sensible response and remains more concrete than most European universities managed. It is an observation that CTU is currently a world-class producer of artificial intelligence and an ordinary consumer of it, and that the second fact is now the more consequential one. Every technical university in Europe can license the same models. Almost none of them have a rector who built an AI research centre, a CIIRC-scale institute, a top-five European computer-vision group, and a national mandate under the Czech AI Strategy to 2030 sitting in the same building as a 51% first-year failure rate. The asymmetry is the opportunity, and it has a short half-life: the differentiation is available for about three years, after which everyone will have done the obvious things.
But attrition is only the first of two problems, and the second changes what the end of a degree is for.
Stanford’s Digital Economy Lab, tracking millions of payroll records, finds employment declining specifically among young workers in the occupations most exposed to AI — entry-level software and technical roles, which is to say the destination of a technical university’s bachelor and master graduates. Set that against the compression findings that make the tutoring case so strong: +34% for novices against near-zero for experts in the NBER field study of 5,179 workers; the largest gains for below-average performers in the BCG field experiment; +9 percentage points for students of the weakest tutors in Stanford’s Tutor CoPilot trial.
Put those together and the shape of the problem is unusual. The technology compresses the performance gap between novice and expert while eroding the jobs in which that gap was historically closed. A CTU graduate can now perform like a competent junior on day one, and may find no junior position in which to become a senior. The apprenticeship function — the first three years in industry where an engineer acquires judgement by being wrong under supervision — is migrating upstream, and the only institution positioned to absorb it is the university.
This is not a distant concern. It bears directly on what a capstone, a diploma thesis, an industrial placement and a laboratory course are for. CESAER’s Engineer of the Future white paper and the CDIO tradition already describe the direction — challenge-based, competence-based, real consequences — without naming this as the reason. Several of the highest-scoring ideas below are that argument made operational.
So the reframe this playbook argues for is not “CTU should adopt AI in teaching”. Everyone will. It is that CTU should treat teaching as an application domain of its own research, and its own students as the population on which the European evidence base gets built — against two problems at once: the third of the cohort lost in year one, and the professional formation that used to happen after graduation and no longer will. That is a claim about identity, not about procurement. A university that publishes on machine learning and runs its education on intuition and committee memory is holding two incompatible epistemologies.
Three conditions make the timing favourable, and all three are transient. The mandate already exists: Strategic Plan CTU 2021+ has four pillars, of which Pillar 1 is Study, and its goals include raising the quality and success rate of study and bringing practice into teaching; Pillar 3 commits CTU to “digitalisation of activities and operations, decision-making on the basis of data”. The evidence has arrived and points at CTU’s worst number, since every well-designed study says the benefit concentrates in the weakest performers. And the university has already run the experiment without noticing — the 2024 annual report records that the Faculty of Mechanical Engineering uses artificial intelligence to predict “at risk” status from weekly examination results and offers targeted help with exam scheduling, credited with a significant fall in failure. It is not in the strategic plan, has no institutional owner, and has never been evaluated to any standard CTU would accept from a doctoral student. The programme does not start from zero; it starts from an unowned success nobody has scaled.
How the ideas are scored
Forty interventions follow. Three dimensions, each out of 10, summed to a composite out of 30. A fourth consideration — feasibility — is deliberately kept out of the score and reported separately, because feasibility should determine sequence, not merit. Scoring hard things down because they are hard is how institutions end up with a portfolio of easy things that changed nothing.
Evidence quality (E) — how well the international library actually supports this. A 9 or 10 means multiple randomised or quasi-experimental studies converging, in comparable settings, on a learning or progression outcome. A 6 or 7 means consistent professional consensus, regulator guidance, or strong observational evidence from named deployments. A 3 or 4 means it is a reasonable inference from adjacent evidence but nobody has tested this thing. A 1 or 2 means it is a bet.
Depth (D) — how structural the change is. Does it alter the teaching production function, what a degree certifies, or who is accountable for what — or does it add a capability on top of an unchanged system? A 9 or 10 changes what CTU is for the student. A 5 or 6 changes how something is delivered. A 2 or 3 is an improvement that leaves every underlying structure intact. High depth is not automatically good; it measures leverage and risk simultaneously.
Expected outcome (O) — magnitude times probability, at CTU specifically. Not “is this a good idea in general” but “what is the realistic expected value of doing this at an institution with 18,168 students, 2,247 academic staff, 221 study programmes across eight faculties and six institutes, a 31.8% first-year failure rate, 19.7% international enrolment, and a rector who is an AI professor”. An idea with a large effect that will probably not survive institutional contact scores lower than a modest effect that certainly will.
Feasibility, reported separately as one of four labels: this year · this cycle (2–3 years) · accreditation-bound (lands at the next programme reaccreditation) · hard (requires external agreement, money that does not exist, or a cultural change with no current sponsor).
The composite is a ranking instrument, not an oracle. Two honest limitations. It rewards ideas whose effects are measurable, which biases against slow cultural moves that matter and cannot be counted. And the E scores are drawn from a library assembled for this question — a different library would move them. Both are reasons to read the reasoning rather than the number.
The scoreboard — forty ideas, ranked
Read as: E evidence quality · D structural depth · O expected outcome at CTU · = composite /30 · feasibility label.
Tier 1 — the seven that clear 24
1. Two-lane assessment at programme level — E 9 · D 9 · O 8 · = 26 · accreditation-bound. Replace MP 5/2023’s activity-by-activity permission schema with a small number of properly secured certification points per programme and everything else open and taught. Scores highest because the evidence is unusually complete on both halves — detection demonstrably fails, and TEQSA’s two-lane model has a five-year institutional track record — and because it changes what a CTU degree asserts, which nothing else on this list does as directly.
2. Gateway Tutor on the five worst first-year courses — E 9 · D 7 · O 9 · = 25 · this cycle. Constrained, course-grounded tutors on the gateway subjects at FS, FD and FJFI, run as randomised trials. Highest expected outcome on the list: four independent RCTs support the intervention, every compression study says the effect concentrates in exactly the population CTU is losing, and the target number is published annually.
3. Discipline-specific verification standards — E 8 · D 9 · O 8 · = 25 · this cycle. Define, per discipline, what counts as checking a machine-produced result — in circuits, in structures, in code, in control. Scores here because the jagged-frontier evidence shows people cannot locate the capability boundary untaught, and because this is the one graduate competence that is both load-bearing and currently taught nowhere.
4. The Junior Engineer Compact — E 7 · D 10 · O 8 · = 25 · hard. With Czech industry: a structured final-year track of supervised responsibility for real work, absorbing the apprenticeship function that entry-level roles no longer perform. The deepest idea on the list and the only one addressing the collapse of the junior pipeline. Marked hard because it requires industry agreement CTU does not yet have.
5. Four graduate AI outcomes in every accredited programme — E 8 · D 9 · O 8 · = 25 · accreditation-bound. Verification, frontier judgement, unassisted core reasoning, accountability for machine-produced results — written into programme documentation with named assessment points. Scores high on depth because accredited outcomes are the only teaching change that survives a change of dean.
6. Teaching redesign counts for promotion and habilitation — E 6 · D 10 · O 8 · = 24 · hard. The single highest-leverage move on the list and the one with the weakest direct evidence, which is why it sits sixth rather than first. Every study of stalled adoption identifies incentives; none of them tested changing incentives. Depth 10 because it is the only idea here that changes what the institution rewards.
7. Assessment twins — E 8 · D 8 · O 8 · = 24 · this cycle. Pair an open, AI-permitted component with a short secured oral or practical assessing the same outcome, scheduled close together for cross-verification. Scores just below two-lane assessment because it is the implementable pattern inside that model rather than the model itself.
Tier 2 — strong (20–23)
8. CS1/CS2 outcome rewrite at FIT and FEL — E 9 · D 7 · O 7 · = 23 · accreditation-bound. Rebuild introductory programming outcomes around decomposition, specification, verification, debugging and reading unfamiliar code. Best-evidenced curriculum change available; capped on outcome only because it touches two faculties.
9. A CTU-built, course-grounded tutoring layer — E 8 · D 8 · O 7 · = 23 · this cycle. Own the pedagogically-constrained layer rather than renting a general assistant. Georgia Tech’s 76.7%-versus-31.3% accuracy gap is the evidence that grounding, not model choice, is what works.
10. The Teaching Evidence Unit — E 9 · D 7 · O 7 · = 23 · this year. Three to four people who make every deployment a pre-registered trial. Highest evidence score on the list; outcome capped because its effect is entirely indirect — it changes the quality of every other decision rather than any student’s result.
11. ETH Zurich lecturer framework, workload-credited, by discipline — E 8 · D 7 · O 7 · = 22 · this year. Adopt rather than draft; deliver in faculty cohorts with hours credited, not added. Scores well and is capped by the honest observation that no study in the library evaluates a faculty AI-development programme against a teaching outcome.
12. The Confusion Map — E 6 · D 8 · O 8 · = 22 · this cycle. Publish, per course, where students actually get stuck, derived from the tutor interaction stream. A university has never before been able to see confusion at scale in the student’s own words; ten million CS50 queries is a map of what is hard about introductory computing that no pedagogical intuition could produce.
13. Retire detection from misconduct procedure — publicly, with the reason — E 9 · D 5 · O 7 · = 21 · this year. Fourteen detectors failed systematic testing; GPT detectors misflag roughly 61% of non-native English writers. With 3,577 international students and 78 English-taught programmes, this is a live equity exposure. Cheapest high-scoring move available.
14. Convert saved lecture hours into studio and seminar contact — E 6 · D 8 · O 7 · = 21 · this cycle. The productive use of every efficiency elsewhere on this list. Scores on depth because it is the difference between an agentic university and an automated one, and it will be decided in budget meetings rather than strategy documents.
15. Universal AI-delivered pre-matriculation bridge — E 7 · D 6 · O 8 · = 21 · this year. Scale what FIT, FEL and FJFI already run in fragments — FIKS, FEL Camp, the Preparatory Week, the Mathematical and Physics Minimum — into a universal, AI-delivered bridge for every admitted student. High outcome because it attacks the failure before enrolment, when it is cheapest.
16. University-wide Defence Week — E 6 · D 8 · O 7 · = 21 · this cycle. A scheduled institutional rhythm of oral authentication rather than per-course vivas invented by exhausted individual lecturers. Turns the staffing problem of secured assessment from a distributed impossibility into a timetabling exercise.
17. Redesign fellowships — E 7 · D 7 · O 7 · = 21 · this cycle. Pay academics released time to rebuild a specific course, with a deliverable. The Ithaka evidence is unambiguous that lack of time, not lack of willingness, is the binding constraint.
18. One governance register with a risk-class sequencing rule — E 8 · D 6 · O 7 · = 21 · this year. AI Act Annex III classification, ESG route, data flows and named owner for every system, plus a published rule that no high-risk deployment precedes an evaluated low-risk one. Protects the programme from its own enthusiasts.
19. Examiner calibration — E 6 · D 7 · O 7 · = 20 · this cycle. Use AI to measure and reduce inter-examiner variance. A larger and more measurable fairness problem than cheating, entirely unaddressed, and newly tractable.
20. Students build the agents — E 4 · D 9 · O 7 · = 20 · this cycle. FIT and FEL students build the tutoring agents for FS and FSv gateway courses as assessed coursework. Collapses the cost, produces an authentic capstone with a real user, and makes the university’s own teaching the object of student engineering. Evidence score is low because nobody has published this; depth is high because it changes who does the work.
21. Retain the interaction stream institutionally — E 6 · D 8 · O 6 · = 20 · this year. MP 5/2023 already states the fact correctly: no AI tool used at CTU is operated by CTU. Today the richest teaching-improvement dataset the university could own accrues to a vendor.
22. The AI-native capstone — E 5 · D 8 · O 7 · = 20 · accreditation-bound. Every capstone ships and publicly defends an artefact built with agentic tooling, assessed on design decisions and verification rather than authorship.
Tier 3 — worth doing (16–19)
23. Diagnostic competence map at entry — E 5 · D 7 · O 7 · = 19 · this cycle. Replace a single admission score with a per-topic gap map that routes the student to specific remediation.
24. “Machines and Judgement” spine course across all eight faculties — E 6 · D 7 · O 6 · = 19 · accreditation-bound. One shared, discipline-adapted course carrying the AI-literacy and verification content, replacing the licensed general online course currently doing that job.
25. The Course Concierge — E 8 · D 4 · O 7 · = 19 · this year. Syllabus and logistics agent. Low depth by design; it is the cheapest archetype, the fastest visible staff relief, and the safest place to learn to run any of this.
26. Teaching-AI as doctoral topics at CIIRC and the AI Center — E 5 · D 7 · O 7 · = 19 · this cycle. Solves staffing, cost and publication simultaneously by making the teaching engine a research programme rather than unpaid service.
27. The Week-Six Trigger — E 6 · D 5 · O 7 · = 18 · this year. Detect and remediate at the specific point in a cumulative course where recoverable falling-behind becomes unrecoverable.
28. Audit and scale the FS at-risk model — E 7 · D 4 · O 7 · = 18 · this year. It already runs, it is credited with a fall in failure, it has no institutional owner, and it has never been audited for cohort drift or subgroup fairness — which the dropout-prediction literature says is where these models fail.
29. Second-chance architecture — E 4 · D 7 · O 7 · = 18 · this cycle. A structured re-entry path with AI-supported catch-up for the near-miss share of the 25.71%, instead of treating failure as terminal.
30. The Removal Register — E 4 · D 8 · O 6 · = 18 · accreditation-bound. Require every programme to name what it retired this cycle. Curricula only ever accrete; adding AI content without a removal instrument produces an unteachable degree.
31. Embedded CIIRC engineer per faculty — E 4 · D 7 · O 7 · = 18 · this cycle. A semester-long residency that transfers capability rather than delivering a system and leaving.
32. Czech technical-language evaluation set — E 5 · D 7 · O 6 · = 18 · this cycle. Frontier models serve Czech technical instruction measurably worse than English. CTU has the NLP capability to build the benchmark, and no one else in the country will.
33. Write teaching-AI into the structural-fund proposals — E 4 · D 6 · O 8 · = 18 · this year. The AIML Research Centre and AI European Centre of Excellence are in preparation now. A workstream written in at drafting stage is funded; one added later is not. Pure timing value, and the window closes.
34. Process portfolio assessment — E 5 · D 7 · O 5 · = 17 · accreditation-bound. Assess the trajectory of work rather than the final artefact.
35. The cognitive gym — E 6 · D 6 · O 5 · = 17 · this cycle. Deliberately unassisted practice spaces, defended pedagogically rather than punitively — the capability on which verification skill is parasitic.
36. EuroTeQ workstream and EDIH industry microcredentials — E 5 · D 6 · O 6 · = 17 · this cycle. The export move. Worthless before delivery exists, valuable immediately after.
37. Validate the admission test against post-AI outcomes — E 5 · D 6 · O 5 · = 16 · this cycle. If AI compresses performance differences, the instrument that used to predict who succeeds may no longer predict it. Nobody has checked.
Tier 4 — scored, and not recommended now (≤15)
38. Assessment variant generation at scale — E 5 · D 4 · O 6 · = 15 · this year. Twenty equivalent exam versions, generated and human-checked. Genuinely useful — but it is a component of the secured lane, not an initiative, and promoting it to a programme invites building the tool before deciding the assessment model it serves.
39. Resit and exam-scheduling optimisation — E 5 · D 3 · O 5 · = 13 · this year. Real efficiency, no structural change, and it risks becoming the visible “AI project” precisely because it is easy and uncontroversial. Do it as operations, not as strategy.
40. Peer observation of AI-mediated teaching — E 5 · D 4 · O 4 · = 13 · this year. Sound practice, but with 2,247 academic staff it consumes exactly the senior attention the redesign fellowships need, and the evidence that observation changes teaching behaviour is weak.
Deliberately not on this list, and why. Automated summative grading, AI admissions triage and remote proctoring were considered and excluded rather than scored, because all three sit inside AI Act Annex III’s high-risk categories while CTU has no compliance track record, no evaluation function and no institutional trust built. They are not bad ideas permanently; they are bad ideas first, and the cost of that mistake is not recoverable.
Tier 1 in full
1. Two-lane assessment at programme level — 26/30
In short. Retire the activity-by-activity permission schema of Methodological Instruction 5/2023 and replace it with a programme-level architecture: a small, deliberately-chosen set of secured certification points where CTU asserts that a named human holds a capability, and everything else open, AI-permitted and taught.
The mechanism, and why it is not a rules change. The current instrument asks, of each activity, may a student use AI for this? That question has no enforceable answer, and the enforcement evidence is conclusive: Weber-Wulff and colleagues, working through the European Network for Academic Integrity, found fourteen detection tools neither accurate nor reliable and defeated by light paraphrase; Liang and colleagues at Stanford found GPT detectors misclassifying roughly 61% of essays by non-native English writers as AI-generated while performing near-perfectly on native speakers. With 3,577 international students — 19.7% of the body — from over 100 nationalities and 78 of 221 programmes taught in English, a detection-founded regime at CTU accuses its international cohort at several times the domestic rate on the basis of second-language fluency.
The two-lane model asks a different and answerable question: what is this assessment for? TEQSA’s 2023 discussion paper established the principles and its 2025 follow-up reports what institutions actually built from them. Assessment of learning certifies — and therefore requires secured conditions and identity assurance, at programme level rather than in every task. Assessment for learning develops — and there AI use is open, expected, and frequently the subject of the assessment. The unwinnable arms race came from demanding both jobs from every assignment.
Why it scores 26. Evidence 9: both halves are unusually well established — the failure of detection empirically, the two-lane model through five years of regulator-guided institutional practice. Depth 9: it changes what a CTU degree asserts to a stranger, which is the institution’s actual product. Outcome 8: high confidence of adoption because it reduces staff burden rather than adding to it, capped only by accreditation timing.
What CTU already has. More than most European universities. MP 5/2023 (ČVUT_MP_2023_05_V01, eight pages, effective 25 September 2023, issued by the Vice-Rector for Bachelor and Master Studies) already did the hard analytical work of thinking activity-by-activity about where AI use is pedagogically load-bearing — that analysis is reusable, it is the form that must change. Two of its provisions should be preserved verbatim: the warning that no AI tool used at CTU is operated by CTU and that all user–tool communication is visible to the operator, and the explicit treatment of deepfake identity modification in online examinations as a disciplinary offence. The Study and Examination Code, consolidated and effective from 1 December 2025, is the harder vehicle in which secured-assessment requirements must ultimately live.
The first move. Run the QAA four-step triage across every programme, through existing internal quality assurance rather than as an emergency parallel process — which is also what keeps it accreditable with NAÚ under the recommended procedures for preparing study programmes. For each programme the output is a single page: which outcomes require certification, at which points, under what identity assurance. Expect the honest answer to be three to five points across a bachelor’s degree, not thirty.
The failure mode. Two, both common. The first is that “secured” is implemented as surveillance — proctoring software, which is both an Annex III high-risk use and, on the evidence in this library on proctoring and disability, an accessibility liability. The second is that the open lane is declared and then quietly policed anyway, with informal suspicion migrating into marking. The countermeasure to both is measurement: publish the misconduct-rate disparity between international and domestic students annually.
Cost and owner. Vice-Rector for Studies, through faculty study committees and the Academic Senate. Drafting is cheap; the real cost is contact hours for secured oral components, which idea 16 (Defence Week) exists to make affordable. The Integrevise research report gives the staffing and cost model for oral assessment at cohort scale — compute it for CTU’s actual cohorts rather than dismissing the option on intuition.
The number that proves it wrong. If, two years in, no programme has reduced its number of assessed tasks and the misconduct disparity has not narrowed, this was a document change and not a reform.
2. Gateway Tutor on the five worst first-year courses — 25/30
In short. Constrained, course-grounded AI tutors on the gateway subjects that fail the most students — mathematical analysis, physics, mechanics, first programming — beginning at the Faculty of Mechanical Engineering, deployed as randomised trials.
The mechanism. First-year failure at a technical university has a stereotyped shape: the material is strictly cumulative, a student falls two weeks behind, week seven becomes unintelligible without week five, the only remediation is a consultation hour that clashes with another lecture, attendance stops, formal failure follows in February. Nothing in that sequence requires a human at the moment of intervention. It requires a patient, correct, course-specific explanation at eleven at night, which is the one thing this technology unambiguously supplies — provided it is built to withhold.
Why it scores 25. Evidence 9: four independent randomised trials converge — Kestin and colleagues’ Harvard physics crossover trial, where students learned more than twice as much in less time than in expert-led active learning; the World Bank’s six-week Nigerian RCT at 0.31 standard deviations; Stanford’s Tutor CoPilot at +4 points overall and +9 points for students of the weakest tutors; and a pooled meta-analytic effect of g = 0.670 across 35 studies and 4,193 participants. Outcome 9, the highest on the list, because the compression findings say the effect lands exactly where CTU’s losses are and the target metric is already published. Depth only 7 — it improves delivery of an unchanged curriculum, which is precisely why it is safe to do first.
What CTU already has. The Faculty of Mechanical Engineering already uses artificial intelligence to predict “at risk” status from weekly examination-period results and offers targeted help with exam scheduling — recorded in the 2024 annual report and credited with a significant fall in failure. CTU’s teaching-AI programme starts from an unowned success nobody has scaled. It also has a substantial existing scaffolding culture to attach to (FIKS, FEL Camp, the FJFI Preparatory Week and its free senior-student tutor system, FD’s mathematics and physics tutoring, the CIPS, ELSA and KC counselling centres), KOS and Moodle as substrate, and in CIIRC and the FEL AI Center the capability to build this properly.
The first move. Two courses, not twenty. Assemble each course corpus — lecture notes, problem sets, worked solutions, past examinations, rights-cleared textbook material — and build retrieval-grounded tutors over them with the pedagogical constraint specified as an engineering requirement: one step at a time, question before answer, never the final result to an assessed problem. Georgia Tech’s numbers are the argument for grounding over model choice: 76.7% answer accuracy against 31.3% for a generic assistant baseline on the same evaluation, and coverage climbing from 21% at 80% precision to over 96% at over 86% precision through iteration.
Then randomise. A waitlist crossover — half the cohort in semester one, half in semester two — is ethically clean, methodologically standard, and produces the single most valuable unpublished result in European educational AI: does a constrained tutor reduce first-year engineering attrition, and by how much. No study in this library answers that.
The failure mode. Instruction dilution, with a published number: CS50 reports 22% of ten million responses containing code blocks despite instructions not to give solutions, 48% at conversation level. Their fix was not a better prompt but teaching fellows reviewing and correcting behaviour. Budget the loop or buy the leak.
Cost and owner. Vice-Rector for Studies with build capability seconded from CIIRC or the FEL AI Center, course ownership retained by the department — never an IT project, which is the documented way this fails. Inference is trivial: CS50 ran at $1.50 per student per year. The real costs are roughly 25 hours of corpus preparation per course and a permanent fraction of a teaching-assistant post per course in supervision.
The number that proves it wrong. First-year bachelor failure at FS, currently 51.1%, against the randomised control. A programme that cannot move it has failed and should be said to have failed.
3. Discipline-specific verification standards — 25/30
In short. Define, write down and assess what it means to check a machine-produced result in each of CTU’s disciplines. Not “critical thinking about AI” as a generic disposition — the specific, technical, discipline-bound question of how an electrical engineer establishes that a circuit analysis is right, how a structural engineer establishes that a load path is right, how a programmer establishes that unfamiliar code does what it claims.
The mechanism, and why this is the load-bearing competence. Two findings define it. Dell’Acqua, Mollick and Lakhani’s field experiment with 758 BCG consultants found that inside the AI’s capability frontier participants produced work rated 40% higher in quality, while on a task just outside it they were 19 percentage points less likely to reach the correct answer than colleagues working with no AI at all — the tool did not merely fail to help, it degraded performance, because the boundary is invisible from inside and the failures are fluent. And Lee and colleagues at Microsoft Research and Carnegie Mellon, surveying 319 knowledge workers across 936 task examples, found that generative AI shifts effort toward verification and integration while higher confidence in the AI predicts less critical-thinking effort — the work moves to a task people are increasingly disinclined to perform.
For an engineer this is not productivity. A structural calculation, a control loop, a dosage algorithm, a safety interlock: being unable to tell that a plausible answer is wrong is how people are harmed. Verification is the competence on which professional liability rests, and it has never been taught explicitly because it used to be a by-product of doing the work by hand.
Why it scores 25. Evidence 8: the jagged-frontier and confidence findings are strong, replicated in shape across settings, and directly on point — though no study has yet taught verification and measured the result. Depth 9: it changes what the degree certifies. Outcome 8: high, because engineering already has the assessment forms to carry it and CTU’s disciplines are exactly the ones where “correct” is externally checkable.
What CTU already has. The single biggest structural advantage in this entire report and the one CTU under-uses: engineering assessment is already checkable against something that is not an examiner’s impression. A bridge calculation is checked by statics. A circuit oscillates or does not. Where humanities faculties must reconstruct authenticity from first principles, engineering mostly has to stop drifting away from it. CTU also has EUR-ACE/ENAEE and ABET-style outcome frameworks as the accreditation vocabulary, and the EuroTeQ Framework of Qualifications as the alliance-level architecture to write this into.
The first move. Commission each faculty to produce a two-page verification standard: the three to five checking procedures a graduate of that discipline must be able to perform, with worked examples of a fluent-but-wrong machine output in that domain and the procedure that catches it. Then assess it directly — give students tasks on both sides of the capability frontier and mark them on whether they knew which was which. That assessment is straightforward to build and almost nobody is building it.
The failure mode. Generic drift. The moment this becomes a university-wide module on “critical evaluation of AI outputs” it is worthless, because verification is not transferable across domains — checking a truss is nothing like checking a compiler optimisation. The countermeasure is that faculties write their own and are not permitted to adopt each other’s.
Cost and owner. Cheap in money, expensive in senior academic attention: this must be written by people who actually verify things professionally. Owner is each faculty’s study dean, coordinated by the Vice-Rector for Studies.
The number that proves it wrong. Every programme has a written verification standard with named assessment points within two accreditation cycles — or this was a memo.
4. The Junior Engineer Compact — 25/30
In short. With Czech industry: convert the final year of CTU’s engineering degrees into a structured track of supervised responsibility for real work with real consequences, formally absorbing the apprenticeship function that entry-level employment used to perform and increasingly does not.
The mechanism, and why this is the deepest idea here. The compression finding that makes idea 2 so strong contains a long-term problem that nobody has solved. AI raises the floor: +34% for novices against near-zero for experts in the NBER study of 5,179 workers; the largest gains for below-average performers at BCG; +9 points for students of the weakest tutors. Simultaneously, Stanford’s Digital Economy Lab finds employment declining specifically among young workers in AI-exposed occupations — precisely the entry-level technical roles where novices historically became experts by being wrong under supervision.
So the technology compresses the novice–expert performance gap while eroding the institution in which that gap was closed. A graduate performs like a competent junior on day one and may find no junior role in which to become a senior. Someone has to run the apprenticeship. The employer’s incentive to do it falls as the productivity gap between a graduate and an experienced engineer narrows. The university is the only remaining candidate.
Why it scores 25 with only E 7. Depth 10 — the maximum on this list — because it redefines what the last two years of an engineering degree are for: less coverage, more supervised responsibility, assessed on judgement under consequence rather than on completion. Outcome 8 on a long horizon. Evidence 7 because the diagnosis is well-evidenced (Stanford’s labour data, the compression studies, CESAER’s and CDIO’s challenge-based direction) while the intervention is not — nobody has run this and measured it.
What CTU already has. Deep industrial relationships and the institutional machinery to formalise them: EDIH CTU and the AI-MATTERS Testing and Experimentation Facility as industry channels; ROBOPROX; 380 partner universities; the existing diploma-thesis and industrial-project traditions, which are the seed of this and are currently assessed as documents rather than as professional performance. CESAER’s Engineer of the Future white paper — written by the association CTU belongs to — already argues for challenge-based learning without naming the collapsing junior pipeline as the reason. And strategic-plan goal 1.3, “bring practice into teaching”, is the existing mandate.
The first move. One faculty, one industrial partner, one cohort of twenty. Define what “supervised responsibility” means as an assessed outcome — the student owns a real deliverable with a real deadline and a real consequence, an industry engineer supervises, a CTU academic certifies the learning. The design question that must be answered first, and honestly, is what the firm gets: with AI compression, a final-year student supervised properly is genuinely productive, which is the argument to make rather than appealing to goodwill.
The failure mode. Two. It becomes an internship scheme — unstructured, unassessed, and indistinguishable from what already exists. Or it becomes free labour, which is an ethical failure and will be recognised as one. The guard against both is that the learning outcome is certified by CTU and the assessment is of judgement, not of output.
Cost and owner. Expensive in coordination, cheap in capital; the vice-rector responsible for cooperation with industry, jointly with a faculty willing to be first. Marked hard because it needs an industry agreement that does not currently exist, and because it will take an accreditation cycle to land properly.
The number that proves it wrong. Graduate employment and, more tellingly, time-to-first-independent-responsibility reported by employers. If graduates of the track are not measurably ahead within three years, the hypothesis was wrong.
5. Four graduate AI outcomes in every accredited programme — 25/30
In short. Compress the available competence frameworks into four assessable outcomes every CTU graduate must demonstrate, and require each of the 221 programmes to show where each is taught and where it is assessed — in the accreditation file, not in a strategy document.
The four, and why exactly these. Verification — can establish by independent means whether a machine-produced result is correct in their discipline (idea 3 supplies the discipline-specific content). Frontier judgement — can tell which side of the capability boundary a task sits on, which the BCG experiment shows people cannot do untaught. Unassisted core reasoning — can perform the discipline’s foundational reasoning without assistance, demonstrated at defined points, because verification is parasitic on it: you cannot check what you could never have derived. Accountability for results one did not personally generate — the professional stance of signing off on machine output, which every engineer already does with finite-element packages and library code, and which is now general.
Why it scores 25. Evidence 8: each outcome traces to specific findings rather than to aspiration. Depth 9: accredited learning outcomes are the only teaching change that survives a change of dean, a budget round, or the departure of the enthusiast who started it. Outcome 8: near-certain to persist once landed, discounted for the accreditation-cycle lag.
What CTU already has. The vocabulary exists and should be adopted rather than invented: UNESCO’s AI Competency Framework for Students (twelve competencies, four aspects, Understand / Apply / Create), the joint OECD–European Commission AI literacy framework, DigComp 2.2’s AI-specific knowledge and attitude examples, Digital Promise’s Understand / Evaluate / Use model, and the AI Literacy Heptagon’s translation into higher-education learning objectives. Structurally, CTU has the EuroTeQ Framework of Qualifications — the alliance deliverable defining what a European engineering graduate must be able to do — which is where these belong rather than in a parallel CTU-only scheme, and the NAÚ methodology for programme design as the accreditation route.
The first move. Draft the four outcomes at university level, then require every programme, at its next reaccreditation, to map them: where taught, where assessed, with what instrument. Do the mapping exercise on three pilot programmes first — one from FIT, one from FS, one from FA — because the honest result will be that most programmes cannot currently point to an assessment for any of the four, and it is better to discover that on three than on 221.
The failure mode. Documentation theatre — the outcomes appear in the file, nothing changes in the room. This is the standard fate of graduate-attribute schemes and it is worth naming in advance. The only reliable countermeasure is the assessment column: an outcome with no named assessment instrument is not an outcome.
Cost and owner. Almost nothing in money; considerable political effort across eight faculties; Vice-Rector for Studies with programme guarantors, using CTU’s EuroTeQ representation to push the same four upward so CTU is defining the alliance standard rather than adopting someone else’s.
The number that proves it wrong. Percentage of accredited programmes with all four outcomes mapped to a named assessment. If it plateaus below half, the outcomes were written at the wrong altitude.
6. Teaching redesign counts for promotion and habilitation — 24/30
In short. Change the criteria for promotion and habilitation so that substantial, evidenced teaching redesign counts as academic achievement — and make the evidence requirement real, so that it means an evaluated redesign rather than a claimed one.
Why this is the highest-leverage idea on the list. Every other proposal in this report is delivered by 2,247 academic staff who currently face an incentive structure that rewards publication and, at the margin, tolerates teaching. Ithaka S+R’s interview study across nineteen universities and its large-N national instructor survey describe the same picture: adoption in isolated pockets, driven by individual enthusiasm, blocked by frictions that have nothing to do with technology — no time, no recognition, no clarity, nobody to ask. The multi-institution barriers study shows the obstacles operating independently at individual, departmental and institutional level, which is the crucial finding: fixing time without fixing recognition changes little, because an academic who redesigns a course has spent a semester on something that will not appear in any file that decides their career.
Every one of the delivery ideas here — the tutors, the assessment redesign, the verification standards, the outcome mapping — is a large uncompensated ask of exactly the people whose promotion depends on something else. This idea is the one that changes the denominator.
Why it scores 24 with only E 6. Depth 10: it changes what the institution rewards, which is the deepest change available to any organisation. Outcome 8: if it lands, everything else gets easier; the discount is for the possibility that it is diluted into a box-ticking criterion. Evidence 6 is the honest constraint and the reason it ranks sixth rather than first — the literature identifies incentives as the binding constraint with great consistency, and not one study in this library tests changing them. That asymmetry should be stated rather than hidden: this is the best-diagnosed, least-tested intervention in the field.
What CTU already has. Strategic-plan Pillar 3 covers human resources, and goal 1.2 commits the university to raising the quality and success rate of study — the mandate exists. The habilitation and professorial-appointment framework is partly national and partly institutional, which bounds how far CTU can move alone; the internal promotion and evaluation criteria are entirely CTU’s.
The first move. Do not open the habilitation question first — it is the slowest and most contested. Start where CTU has unilateral control: internal performance evaluation, faculty-level promotion criteria, and the allocation of teaching-relief and institutional support. Define a recognised category of evidenced teaching redesign, with a specific evidentiary bar — a redesigned course, a pre-registered evaluation, a result, published. That bar is what stops the criterion degrading into “attended a workshop”, and it dovetails exactly with idea 10, the Teaching Evidence Unit, which supplies the evaluations.
The failure mode. Dilution into a checkbox, which is how most teaching-recognition schemes end. The countermeasure is the evidentiary bar and the fact that a published evaluation is externally legible in a way that a self-reported innovation is not.
Cost and owner. No direct cost, high political cost; the rector and the Academic Senate. This is the item that most requires the incoming rectorate’s authority, and it is worth noting that a rector who is an AI professor has unusual standing to argue that teaching with these systems is a serious intellectual activity rather than a service task.
The number that proves it wrong. The count of promotions in which evidenced teaching redesign was a material factor. If it is zero after two cycles, the criterion exists on paper only.
7. Assessment twins — 24/30
In short. For each outcome that must be certified, run two deliberately linked components: an open, AI-permitted piece of substantial work, paired with a short secured oral defence or practical demonstration of the same outcome, scheduled close together so each cross-verifies the other.
The mechanism. This is the implementable pattern inside idea 1’s architecture, and it resolves the practical objection that kills most two-lane implementations — that secured assessment at cohort scale is unaffordable. The insight is that the secured component does not have to carry the content; it only has to carry the authentication. A student who has genuinely done the open work can defend it in eight minutes. A student who has not, cannot, and no detector is required to establish that. The paper in this library builds the case through Messick’s validity framework: it shows precisely which assessment types generative AI undermines and why, then proposes the paired-component design as the response.
Why it scores 24. Evidence 8: the validity analysis is rigorous, the oral-assessment operating evidence is solid, and the design is a direct consequence of the two-lane principles that carry regulator backing. Depth 8: it changes the unit of assessment from the artefact to the artefact-plus-defence. Outcome 8: high confidence, because engineering already runs defences for theses and this generalises an existing form rather than importing a foreign one.
What CTU already has. The diploma-thesis defence is exactly this instrument, already operating at scale, already accepted culturally, already in the Study and Examination Code. The move is to generalise a form CTU already trusts down into coursework, not to invent one. The design studio’s crit is the same instrument in another register — and it is worth noticing that the Faculty of Architecture has by far the lowest first-year failure rate at 9.5%, in a faculty where the crit is relentless and work is defended continuously rather than submitted.
The first move. Pick the three or four courses per programme that carry the most certification weight and twin them. Then solve the timetabling, which is the actual constraint and is what idea 16 (a university-wide Defence Week) addresses: eight minutes per student per twinned assessment is impossible when every lecturer schedules it individually and entirely possible as an institutional rhythm. Compute CTU’s real numbers from the Integrevise oral-assessment cost model rather than dismissing it — for a 200-student cohort, one twinned assessment is roughly 27 examiner-hours, which is a scheduling problem, not an impossibility.
The failure mode. Inter-examiner variance. Oral assessment is only fair if examiners are calibrated, and they are typically not — which is why idea 19 sits adjacent to this one and should be done alongside it rather than after. The second failure mode is the defence degrading into a formality that everyone passes, at which point it certifies nothing and costs real hours.
Cost and owner. Contact hours, honestly stated in the workload model rather than absorbed silently by teaching staff — the single most common way assessment reform is quietly sabotaged. Faculty study deans, coordinated centrally for timetabling.
The number that proves it wrong. Pass-rate divergence between the open and secured components. If they agree perfectly, the secured component is not authenticating anything; if they diverge wildly, the open lane has a problem worth knowing about. Either result is informative, which is why this is worth instrumenting from the first cohort.
Tier 2 — the strong middle
These fifteen carry real weight and several are preconditions for Tier 1. Each gets the mechanism, the first move, and the catch.
8. CS1/CS2 outcome rewrite at FIT and FEL — 23. Mechanism: Becker and colleagues’ argument in Programming Is Hard — Or At Least It Used To Be is that many introductory objectives were proxies — the discipline never wanted students to write a loop from memory, it wanted decomposition, and the loop was how it checked. The proxy broke; the objective did not. Prather and colleagues’ observation of CS1 students using Copilot names the new pathologies precisely: drift, where the student’s mental model silently diverges from the accumulating code; over-trust of plausible output; collapse of the metacognitive loop. Ma, Chen and Konomi’s dialogue-log study adds that interaction pattern, not volume, predicts performance — and patterns are teachable. First move: rewrite outcomes around decomposition, specification, verification, debugging and reading unfamiliar code; adopt the ASEE pattern of requiring students to submit their own solution alongside the AI’s and account for the difference. The catch: FIT has 2,456 students and only six programmes — unusually redesignable — but assessment cost rises, because code critique cannot be autograded.
9. A CTU-built, course-grounded tutoring layer — 23. Mechanism: the pedagogy lives in the wrapper, not the model, and grounding is what produces accuracy — Georgia Tech’s 76.7% against 31.3% for a generic baseline. First move: define three explicit layers with owners — Microsoft Copilot under @cvut.cz accounts as the general floor (keep it; the Enterprise Data Protection posture is the hard part and is already solved), a CTU-built course-grounded tutor above it, and an institutional data layer above that. The catch: this only works if CIIRC and the FEL AI Center are commissioned as a research programme with publications and doctoral topics, not conscripted as an internal service desk.
10. The Teaching Evidence Unit — 23. Mechanism: the What Works Clearinghouse handbook and the EEF evaluator guide define what counts as evidence in education, and most published claims about AI in higher education would not qualify — they are satisfaction surveys without comparison conditions. First move: three to four people, reporting to academic governance rather than to the AI programme, with one non-negotiable rule — no deployment goes live without a comparison condition and a pre-registered outcome. The catch: it must not report to the people whose project it evaluates, or it produces evidence theatre. Its first task should be auditing the existing FS at-risk model, not blessing a new build.
11. ETH Zurich lecturer framework, workload-credited, by discipline — 22. Mechanism: the instruments already exist and are unused — UNESCO’s fifteen teacher competencies across five dimensions, DigCompEdu’s twenty-two competences on a six-level ladder, and ETH Zurich’s nine-page AI Competence Framework for Lecturers, built by a peer technical university for exactly this population. First move: adopt and translate ETH’s; map to DigCompEdu for European legibility; deliver one cohort per faculty, taught by academics of that faculty using that faculty’s own courses, hours credited against load. The catch: a mechanical engineer will not attend a generic session on prompt writing, and should not be asked to. Generic delivery is how this fails.
12. The Confusion Map — 22. Mechanism: the tutor interaction stream is a record of confusion at scale, in the student’s own words, at the moment it occurs — not inferred from exam performance months later. Ten million CS50 queries constitute a map of what is genuinely hard about introductory computing that no amount of pedagogical intuition could produce. First move: cluster and publish, per course, the top twenty points of difficulty each semester, to the course team first and the faculty second. The catch: it must be framed as course diagnostics, never as lecturer evaluation, or the data will be resisted and then gamed.
13. Retire detection from misconduct procedure, publicly — 21. Mechanism: fourteen detectors failed systematic testing; GPT detectors misflag roughly 61% of non-native English writers. First move: one paragraph in a rectoral instruction, withdrawing detector output as a basis for proceedings and stating the equity reason, so staff stop relying on it informally. The catch: announce it with the assessment triage from idea 1, or it reads as surrender rather than as reform.
14. Convert saved lecture hours into studio and seminar contact — 21. Mechanism: Georgia Tech’s 500-plus saved teacher hours are real, and they are overwhelmingly logistics and repeated questions. Whether that recovery becomes better teaching or becomes a staffing cut is the entire difference between an agentic university and an automated one. First move: commit in advance, in writing, that recovered hours are redeployed to contact and supervision — before the savings materialise and the budget round arrives. The catch: the supervision loop that makes the tutors work is exactly the line item that gets cut, and nothing visible breaks for several months.
15. Universal AI-delivered pre-matriculation bridge — 21. Mechanism: attack the failure before enrolment, when it is cheapest, and when the compression evidence says the marginal student gains most. First move: scale what already exists in fragments — FIKS at FIT, FEL Camp and the Embedded Technology Club, FJFI’s Preparatory Week plus its Mathematical and Physics Minimum, FSv’s mock entrance examinations, FD’s mathematics and physics tutoring — into one AI-delivered bridge offered to every admitted student, with a diagnostic front end. The catch: uptake is voluntary and self-selects for the students who need it least; the design problem is reach, not content.
16. University-wide Defence Week — 21. Mechanism: secured oral assessment is impossible when every lecturer schedules it alone and entirely possible as an institutional rhythm. First move: two fixed weeks in the academic calendar in which twinned assessments are examined, centrally timetabled, with examiner pools drawn across departments. The catch: it collides with everything else in the calendar and requires the Academic Senate to defend the slot against encroachment for at least three years before it becomes normal.
17. Redesign fellowships — 21. Mechanism: Ithaka’s evidence is unambiguous that time, not willingness, is the binding constraint. First move: twenty fellowships a year, one semester of teaching relief each, with a defined deliverable — a redesigned course, a documented assessment change, and an evaluation. The catch: without idea 6’s promotion criteria this rewards the already-committed and does not change the population; the two should be introduced together.
18. One governance register with a risk-class sequencing rule — 21. Mechanism: Annex III of the AI Act classifies as high-risk exactly the uses an efficiency-minded administration would automate first — admission, evaluation of learning outcomes, level assignment, examination monitoring — while tutoring, explanation and formative feedback are not on the list. First move: one register listing every AI system in teaching with its Annex III classification, provider-or-deployer status, data flows, human oversight, QA route and named owner; plus a published rule that no high-risk deployment precedes an evaluated low-risk one. The catch: modest work before the pilots, a project after the third.
19. Examiner calibration — 20. Mechanism: inter-examiner variance is a larger, more measurable and more consequential fairness problem than cheating, and it becomes critical the moment idea 7 puts more weight on oral assessment. First move: double-mark a sample, measure the variance, and use AI to generate calibration exemplars and to flag outlier marking patterns for human review — never to mark. The catch: it is politically delicate in a way detection never was, because the subject of measurement is academic staff rather than students.
20. Students build the agents — 20. Mechanism: FIT and FEL students build the tutoring agents for FS and FSv gateway courses as assessed coursework, supervised by CIIRC. Cost collapses, the capstone becomes authentic with a real user and a real evaluation, and the university’s own teaching becomes the object of student engineering. First move: one project cohort, one course, with the domain academic from the receiving faculty as the client. The catch: evidence 4 — nobody has published this, quality control is a genuine risk, and student-built systems must pass the same evaluation bar as any other before touching a real cohort.
21. Retain the interaction stream institutionally — 20. Mechanism: MP 5/2023 states the fact correctly — no AI tool used at CTU is operated by CTU, and all user–tool communication is visible to the operator. Today the richest teaching-improvement dataset the university could own accrues to a vendor. First move: make institutional retention of interaction logs a requirement of the layer-two build, under the Jisc Code of Practice for Learning Analytics. The catch: it only becomes an asset if someone is funded to read it, which is idea 12.
22. The AI-native capstone — 20. Mechanism: every capstone ships and publicly defends an artefact built with agentic tooling, assessed on design decisions, verification and defence rather than authorship — the CDIO project-based rubric pattern with explicit AI-use criteria. First move: rewrite the capstone rubric in one programme and run it for a year before generalising. The catch: team projects can be substantially machine-generated in ways a busy assessor will not detect, so the defence carries the assessment weight, not the artefact.
Tier 3 — worth doing, in the slipstream of something bigger
These fifteen are real but should not carry a programme. Most become cheap once a Tier 1 or Tier 2 item has built the capability they depend on.
23. Diagnostic competence map at entry — 19. Replace a single admission score with a per-topic gap map generated from a diagnostic, routing each admitted student to specific remediation. It is the front end that makes the pre-matriculation bridge (15) land on the students who need it rather than on the ones who volunteer. Cheap once the bridge exists; pointless before it.
24. “Machines and Judgement” spine course — 19. One shared, discipline-adapted course carrying AI literacy, verification and professional accountability across all eight faculties — replacing the licensed general online course (CTUPRGEAI, “Elements of AI”) currently standing in for this. Scores moderately because a course is a weaker instrument than an outcome: idea 5 does the durable work, and this is where it gets taught.
25. The Course Concierge — 19. A syllabus-and-logistics agent, built on the SyllabusQA pattern with an explicit factuality metric. Depth 4 by design and that is the point: it is the cheapest archetype, gives the fastest visible relief to teaching staff, and lets the institution learn to run one of these where the cost of an error is a corrected deadline rather than a corrupted understanding of thermodynamics. Do it first in time, not first in importance.
26. Teaching-AI as doctoral topics at CIIRC and the AI Center — 19. Makes the engine a research programme rather than unpaid service: doctoral students build and evaluate the tutors, publish the results, and the cost, the staffing and the evidence problem are solved by the same move. Depends entirely on the receiving academics treating it as real work, which is idea 6 again.
27. The Week-Six Trigger — 18. Cumulative courses have a specific point at which recoverable falling-behind becomes unrecoverable, and it is identifiable per course from historical data. Detect it and intervene there rather than at the examination. A refinement of idea 2 rather than a separate programme, and it needs the interaction stream to be worth much.
28. Audit and scale the FS at-risk model — 18. It already runs, the annual report credits it with a significant fall in failure, it has no institutional owner, and nobody has checked it for cohort drift or subgroup fairness — which is precisely where the dropout-prediction literature says these models fail, and fail worst for the groups an institution most wants to help. Low depth, high urgency: this is an unaudited model running on real students today.
29. Second-chance architecture — 18. Total study failure was 25.71% in 2024, and a large share of that is near-miss rather than incapable. A structured re-entry path with AI-supported catch-up converts some of it back. Scores modestly on evidence because nobody has evaluated such a path with AI support, and high on depth because it changes whether failure is terminal.
30. The Removal Register — 18. Require every programme, at each revision, to name what it retired. Curricula only accrete; adding AI competence to 221 programmes without a removal instrument produces a degree nobody can complete. Costs nothing, is universally resisted, and is the quiet precondition for ideas 5, 17 and 24 not making things worse.
31. Embedded CIIRC engineer per faculty — 18. A semester-long residency inside a faculty rather than a system delivered and abandoned — capability transfer, not procurement. Cheap, and it addresses the departmental level of the three-level adoption barrier that neither training (individual) nor policy (institutional) reaches.
32. Czech technical-language evaluation set — 18. Frontier models serve Czech technical instruction measurably worse than English, and no one has quantified how much worse in mechanics, structures or circuits. CTU has the NLP capability, the sovereign interest is real, and 143 of its 221 programmes are not in English. A small, publishable, genuinely national asset.
33. Write teaching-AI into the structural-fund proposals — 18. The AIML Research Centre and the AI European Centre of Excellence are in preparation. A workstream written in at drafting stage is funded from structural funds; the same workstream added afterwards comes out of the teaching budget. Pure timing value, zero intellectual content, and the window is closing now — which is why it appears in the first ninety days of the sequence despite scoring 18.
34. Process portfolio assessment — 17. Assess the trajectory of work rather than the final artefact. Pedagogically strong and administratively heavy; it is the right instrument for design and thesis work and the wrong one for a 400-student service course.
35. The cognitive gym — 17. Deliberately unassisted practice spaces, defended pedagogically rather than punitively — the capability on which verification skill depends, since you cannot check what you could never have derived. Scores modestly because it is a framing of the secured lane rather than an independent intervention, and because it is easily caricatured as nostalgia.
36. EuroTeQ workstream and EDIH industry microcredentials — 17. The export move: propose the teaching-AI programme as a EuroTeQ deliverable across an alliance of 115,000 students, and package the faculty-development programme as a paid microcredential for Czech industry through EDIH CTU. Worthless before delivery exists — announcing leadership before building capability is the most reliable way to discredit the whole programme — and valuable immediately afterwards.
37. Validate the admission test against post-AI outcomes — 16. If AI compresses performance differences, the instrument that historically predicted who succeeds at CTU may have quietly stopped predicting it. Nobody has checked, the data exists, and the answer matters for every recruitment decision. It ranks low only because it is a study rather than an intervention — but it is a cheap study that the Evidence Unit could run in its first year.
The portfolio view — what depends on what
A ranked list is not a plan, because several high scorers are useless without a low scorer underneath them. Four dependency chains govern the sequence.
The assessment chain. Retiring detection (13) is the precondition for the two-lane triage (1), which defines the certification points that assessment twins (7) implement, which are only affordable if Defence Week (16) exists and only fair if examiner calibration (19) runs alongside. Four of those five are Tier 2 or lower, and the Tier 1 item at the top of the chain cannot land without them. This is the clearest case in the report of ranking and sequencing pulling apart.
The delivery chain. The Course Concierge (25) builds the corpus pipeline, the logging architecture and the institutional confidence that the Gateway Tutor (2) then needs. The tutor produces the interaction stream (21), which becomes the Confusion Map (12), which is what makes the whole thing improve rather than merely run. Skipping the Concierge to get to the tutor faster is the commonest and most expensive shortcut available.
The capability chain. Nothing in the delivery or assessment chains is executed by anyone other than the 2,247 academic staff. The lecturer framework (11) supplies the competence, the redesign fellowships (17) supply the time, and the promotion criteria (6) supply the reason. All three are required; any two of them produce a programme that runs on volunteers and stops when they tire.
The evidence chain. The Evidence Unit (10) is upstream of everything, because it is what converts activity into knowledge — and it is also what makes the promotion criteria (6) meaningful, since “evidenced teaching redesign” requires someone to produce the evidence. Standing it up late means the first year’s deployments are unevaluable forever; once every student has the tutor, the counterfactual is gone.
The practical consequence: the first ninety days should be dominated by Tier 2 and Tier 3 items, and that is not a compromise. It is what the dependency structure says.
Thirty-six months
Days 1–90 — the free moves and the closing windows. Retire detection publicly with the equity reason (13). Name the accountable vice-rector. Stand up the Evidence Unit (10) and build the AI Act register (18). Write the teaching-AI workstream into the AIML Centre and AI European Centre of Excellence proposals (33) — this is the only item on the list with a hard external deadline. Commission the Course Concierge (25) and begin corpus preparation for two gateway courses. Audit the FS at-risk model (28), because it is running unaudited today. Adopt and translate the ETH Zurich lecturer framework (11).
Months 4–12 — first delivery and the incentive question. Gateway Tutor version one live in one course as a randomised waitlist trial (2). First discipline cohort of the lecturer programme at FS — the faculty with both the worst failure rate and the strongest incentive to move. Two-lane triage completed for FIT and FEL (1). First twenty redesign fellowships awarded (17). Open the promotion-criteria question (6) — it takes the longest and should start now, at internal evaluation and faculty promotion level where CTU has unilateral control, not at habilitation. Faculties commissioned to draft verification standards (3).
Months 12–24 — depth. Second and third gateway courses; first trial published. Assessment twins piloted in three or four high-certification courses (7), with the first Defence Week timetabled (16) and examiner calibration running alongside (19). Four graduate outcomes drafted and mapped on three pilot programmes (5). CS1/CS2 rewrite at FIT enters accreditation (8). Confusion Map published to course teams (12). Junior Engineer Compact scoped with one faculty and one industrial partner (4).
Months 24–36 — structure. Graduate outcomes into the accreditation pipeline across faculties. Verification standards written and assessed. Junior Engineer Compact running with its first cohort of twenty. Promotion criteria in force. EuroTeQ workstream proposed and industry microcredentials launched (36) — last, because the export is only credible once the delivery exists.
The seven numbers
Everything above reduces to seven indicators, five of which CTU already publishes or could publish tomorrow in “ČVUT v číslech”.
First-year bachelor study failure, by faculty — 31.8% university-wide, 51.1% at FS, 48.1% at FD, 9.5% at FA. The primary outcome. Every other number here is instrumental to this one.
Misconduct-rate disparity, international versus domestic students — measurable today, almost certainly uncomfortable, and the honest test of whether the assessment reform is real or cosmetic.
Programmes with documented secured certification points — currently zero of 221.
Academic staff at a defined DigCompEdu-mapped level, by faculty — currently unmeasured, which is itself the finding.
Trials completed with a control condition and a pre-registered outcome — currently zero. This is the number that decides whether CTU cites others’ evidence for the next decade or generates it.
Promotions in which evidenced teaching redesign was a material factor — currently zero, and the single best indicator of whether anything structural changed.
Time to first independent professional responsibility, reported by employers — the long-horizon test of the Junior Engineer Compact, and the only one that measures whether the degree still does what it claims.
Close — what the scoring actually showed
Three things fell out of scoring forty ideas that did not fall out of choosing ten.
The highest-leverage idea is the worst-evidenced one. Changing what promotion rewards (6) scored a 10 on depth and a 6 on evidence, because the literature identifies incentives as the binding constraint with remarkable consistency and not one study has tested changing them. That asymmetry is worth sitting with. It means the field’s most confident diagnosis is also its least tested prescription, and an institution that acted on it would be doing something genuinely novel rather than following practice. For a university that already publishes on machine learning, running that experiment properly — and measuring it — is a more interesting contribution than another tutoring pilot.
Ranking and sequencing came apart, and the sequence should win. The dependency chains put Tier 2 and Tier 3 work in the first ninety days and leave the highest-scoring item, two-lane assessment, waiting on an accreditation cycle. An institution that executes strictly in score order will stall, because it will attempt the deep structural moves before the cheap enabling ones exist. The score says what matters; the chain says what is possible next.
The second problem is as large as the first, and far less worked on. Attrition is the largest measurable loss and idea 2 is the most reliable way to attack it — that much is settled by the evidence. But the collapse of entry-level technical employment puts an equally large question over the end of the degree, and unlike attrition it has no established answer anywhere. Only one idea in forty attempts one, and it scored a 10 on depth for exactly that reason. A list of forty options that contains a single response to a problem this size is itself a finding about the state of the field.
Which leaves the reframe the whole exercise keeps returning to: CTU does not need to become a different kind of institution to do any of this. It needs to apply to its own teaching the standard of evidence it already demands of its doctoral students — and to reward the people who do it. The first is a method the university already owns. The second is a decision only the rectorate can make, and on the evidence assembled here it is the one that decides whether the other thirty-nine ideas are a programme or a document.




