Written by the ENSI Foresight Division on a library of 107 primary documents — randomised trials, meta-analyses, regulator guidance, institutional reports and deployed-system evaluations — downloaded and indexed in the “AI for Teaching at CTU” library, and cross-read against ENSI’s Education for the Agentic Age library.
Every university in Europe is currently running the same conversation, and it is the wrong one. The conversation is about permission: what may students use, what must they declare, how do we catch them. It produces policy documents, disclosure forms and detection subscriptions, and it produces almost no change in what is actually taught or how anyone actually learns. Meanwhile the evidence base — which has grown from nothing to several hundred studies in three years, and which this library assembles — has been quietly answering a different and far more consequential question: what does a machine that can explain anything, at any hour, at negligible cost, do to the production of an engineer?
The answer that emerges is not the one either camp expected. The optimists expected a tutor in every pocket and a two-sigma revolution; the pessimists expected a generation that cannot think. The measured reality is sharper and more useful than either. Where AI tutoring has been put through a proper randomised trial with a pedagogically constrained system — Kestin and colleagues’ Harvard physics experiment, published in Nature Scientific Reports, where students learned more than twice as much in less time than in an expert-taught active-learning class; the World Bank’s six-week Nigerian trial returning 0.31 standard deviations, among the most cost-effective learning interventions ever measured; Stanford’s Tutor CoPilot RCT across 900 tutors — the effects are large, real, and replicated across wildly different contexts. And where AI has simply been made available as a general-purpose assistant with no pedagogical constraint, the measured effect on thinking runs the other way: Microsoft Research and Carnegie Mellon’s study of 319 knowledge workers found that higher confidence in the AI predicts less critical-thinking effort, with users shifting from producing judgement to verifying someone else’s.
Those two findings are not in tension. They are the same finding stated twice. The variable that decides whether AI raises or lowers human capability is not the model — it is the structure wrapped around the model. Kestin’s tutor was explicitly forbidden to give answers; it was built to withhold, to question, to force retrieval. CS50’s duck at Harvard, the largest educational agent deployment on record at 211,000 students and 10 million queries, was engineered with the same intention — and its own published evaluation reports that 22% of responses contained code blocks despite instructions not to hand out solutions, rising to 48% at conversation level. That is the whole discipline in one number: the pedagogy is in the guardrail, and the guardrail leaks. An institution that buys models and skips the structure has bought the failure mode without the benefit.
This is why “should we allow it” is the wrong conversation for a technical university in particular. The prevalence question is settled: the HEPI/Kortext survey of UK students found 92% using generative AI, up from 66% a year earlier, and 88% using it in assessed work. MIT’s own institution-wide survey, reported by its Ad Hoc Committee in August 2026, found 46% of undergraduates using LLMs daily — and, more damningly, that while two-thirds of students believed AI would matter in their careers, only 25% felt MIT was adequately preparing them to use it. That gap is the actual institutional failure. It is not a discipline problem. It is a curriculum problem, an assessment problem, and above all a staff-capability problem, and every month spent litigating permission is a month not spent closing it.
There is a further reason a technical university cannot treat this as a general higher-education issue with a general higher-education answer. The disruption is not uniform across disciplines; it is concentrated almost exactly where a technical university lives. Computing education has the deepest evidence base and the sharpest dislocation — the ITiCSE working group’s landmark survey, Becker and colleagues’ “Programming Is Hard — Or At Least It Used To Be”, and the GitHub/Microsoft/MIT randomised trial in which Copilot users completed a programming task 56% faster — because code is the modality large language models are best at. The graduate attribute that a Faculty of Information Technology has spent decades certifying is precisely the one whose market price is moving fastest. Meanwhile Stanford’s Digital Economy Lab, tracking millions of payroll records, finds employment declining specifically for young workers in the occupations most exposed to AI — the entry-level software and technical roles that are the destination of a technical university’s undergraduates. The institution’s product is being repriced from both ends simultaneously.
And yet the same discipline structure that creates the exposure also supplies the defence, which is the most under-appreciated point in this entire literature. Engineering education already possesses, as its native pedagogical form, the thing every other faculty is now scrambling to invent: assessment against a physical or functional artefact that must actually work. A bridge calculation is checked by statics, not by an examiner’s impression of the prose. A circuit either oscillates or it does not. A robot either completes the task or falls over. CDIO, the international engineering-education framework to which the project-based assessment literature in this library belongs, has been building this for twenty-five years; the CESAER white paper on the engineer of the future, written by the very association of European technical universities that a school like ČVUT belongs to, argues for exactly this competence-and-challenge-based direction. Where the humanities must now reconstruct authenticity from first principles, engineering mostly has to stop drifting away from it — away from the worksheet, the boilerplate lab report, the individually-submitted problem set that a model completes in nine seconds.
So the reframe that organises this report is this. AI in teaching is not a tool question, it is a capability question — and the capability being built or destroyed is the institution’s, not the student’s. Universities that treat generative AI as software to be procured and policed will get the measured downside: offloaded thinking, unenforceable rules, an integrity arms race they lose, and graduates who use models badly because nobody taught them to use models well. Universities that treat it as a redesign of the teaching production function — what is assessed, what is taught, what staff can do, what data the institution keeps and what it is allowed to keep — get the measured upside, which is genuinely large, and get it disproportionately for their weakest students, which is where a technical university’s real losses are concentrated. The evidence for that last claim is the most decision-relevant thing in this library, and it is finding number two.
The findings in brief
The tutoring effect is real, large and replicated — but only for systems built to withhold. Harvard’s constrained physics tutor beat expert-led active learning by more than 2×; Nigeria returned 0.31 SD; the pooled meta-analytic effect of ChatGPT on learning is g = 0.670. Unconstrained chatbots do not reproduce it.
The gains concentrate in the weakest performers. Novices gained +34% in the NBER field study versus almost nothing for experts; Tutor CoPilot helped students of lower-rated tutors most. For an institution with 30%+ first-year failure, this is the single highest-value finding in the library.
AI-text detection does not work, and building policy on it is an equity liability. Fourteen detectors failed systematic testing; GPT detectors flagged ~61% of non-native-speaker essays as AI-written. Assessment redesign is the only durable lane.
Computing and programming education is the most disrupted subject on earth and has the best evidence about it — and the disruption is to learning objectives, not merely to cheating.
Over-reliance is measurable, and it is the real cost — not plagiarism. Confidence in AI predicts reduced critical-thinking effort; students know it and are frightened of it (90% of MIT respondents concerned about overreliance).
A deployed teaching agent is astonishingly cheap and demonstrably useful — and it leaks. CS50: $1.50 per student per year, 94% found it helpful, 22% of responses handed out code anyway.
Faculty capability, not technology, is the binding constraint — and the frameworks to fix it (UNESCO, DigCompEdu, ETH Zurich’s lecturer framework) already exist and are unused.
Engineering’s artefact-based pedagogy is a structural advantage nobody is exploiting — the lab, the design studio and the capstone are already AI-resistant assessment.
The law has already decided most of this is high-risk. The EU AI Act’s Annex III names admission, outcome evaluation, level assignment and exam monitoring as high-risk uses — which is most of what a university would want to automate.
Almost none of it has been evaluated to the standard the university would demand of its own research — and early-warning models specifically degrade across cohorts and subgroups.
How this report is organised
The ten findings are ranked by strength of evidence and decision-weight combined — how well the claim is established, and how much changes if you act on it. Four tests decide the ranking: whether the core result rests on randomised or quasi-experimental designs rather than surveys of opinion; whether independent teams in independent contexts converge on it; whether the outcome measured is learning or performance rather than satisfaction; and whether anything has been replicated at scale in a real institution rather than a laboratory. A result from a pre-registered RCT with a learning outcome outranks a beautifully-argued position paper. A deployment evaluation with published failure rates outranks a vendor case study. Sector surveys of what administrators intend rank last, however often they are quoted.
Each finding follows the same discipline: the claim in one line; the actual studies from the library that support it; an explicit statement of what the evidence does not show, because this field is drowning in over-claiming in both directions; and the institutional move it obliges. Where a finding bears specifically on a mid-sized European public technical university — the case that motivated this library — that is drawn out at the end of the section rather than allowed to colour the evidence. The companion report, The ČVUT Playbook, does the institution-specific work in full.
1. The tutoring effect is real and large — but only for systems built to withhold
Where AI tutoring has been tested properly, the effect sizes are among the largest in the history of education research. Every one of those trials used a system deliberately engineered not to answer.
Start with the strongest single study in the library. Kestin, Miller and colleagues ran a randomised crossover trial in Harvard’s PS2 physics course, published in Nature Scientific Reports in 2025. Students learned the same material either in a class taught by expert instructors using active-learning methods — already the gold standard against which everything else underperforms — or with an AI tutor built on GPT-4 and constrained by a detailed pedagogical prompt. Students in the AI condition learned more than twice as much, in less time, and reported higher engagement and motivation. The comparison is what makes it remarkable: this is not AI versus a bad lecture. It is AI versus the best-evidenced form of human undergraduate physics teaching, in one of the world’s most selective institutions, and the AI condition won on a pre-specified learning measure.
The result does not stand alone. The World Bank’s From Chalkboards to Chatbots working paper reports a six-week randomised trial of GPT-4-based virtual tutoring in Nigeria returning 0.31 standard deviations — which the authors, using standard learning-adjusted year conversions, place at the equivalent of one and a half to two years of business-as-usual schooling, and among the most cost-effective interventions ever measured in the education literature. Stanford’s Tutor CoPilot trial — the first RCT of a human–AI tutoring system in live tutoring, across 900 tutors and 1,800 students — found +4 percentage points of topic mastery overall at roughly $20 per tutor per year. An exploratory randomised trial in authentic UK classrooms, included here as the closest European deployment evidence, tested both learning effect and safeguarding simultaneously. And the meta-analytic anchor pools 35 experimental studies across 4,193 participants and reports g = 0.670 for ChatGPT’s effect on learning outcomes, with no significant publication bias detected.
For context on how good that is, the library carries both of the necessary baselines. Bloom’s 1984 paper — the source of the two-sigma claim that every AI-tutoring pitch deck implicitly invokes — established one-to-one human tutoring plus mastery learning as the ceiling nobody could afford. The pre-LLM reality check is the Ma, Adesope, Nesbit and Liu meta-analysis in the Journal of Educational Psychology: 107 effect sizes across 14,321 learners found intelligent tutoring systems outperforming large-group instruction and non-ITS computer instruction, but not outperforming individualised human tutoring or small-group instruction. Thirty years of intelligent tutoring systems bought a real but modest gain at high engineering cost per subject. The current generation is producing larger effects, in weeks, in domains nobody hand-authored.
Now the constraint that makes the whole finding conditional. In every trial above, the system was built to withhold. Kestin’s tutor was instructed to reveal one step at a time, to ask before telling, to refuse to produce the final answer. Tutor CoPilot does not tutor the student at all — it coaches the human tutor in real time. CS50’s duck at Harvard, whose own published evaluation is the most honest deployment document in this library, was likewise engineered to refuse solutions. There is no trial in this library in which handing students an unconstrained frontier model and letting them chat produced a learning gain of this magnitude. The pedagogy is not in the model. It is in the wrapper — the system prompt, the retrieval over the actual course material, the refusal policy, the scaffolding sequence.
What the evidence does not show. It does not show that AI tutors beat good teaching in general — Kestin’s design gave the AI condition a highly structured, expertly-prompted tutor in a single well-defined physics topic, which is a favourable case and was designed to be. It does not show durability: almost every study measures learning at or near the end of the intervention, and the library contains no strong evidence on retention months later. It does not show that gains survive when the tutor is built by a busy academic rather than a research team. And the meta-analytic g = 0.670 pools studies of highly variable quality with mostly short interventions and mostly proximal outcome measures — the standard conditions under which education effect sizes later shrink.
What it obliges. Stop procuring chatbots and start commissioning constrained tutors, subject by subject, grounded in the institution’s own course material, with the refusal behaviour specified as a pedagogical requirement rather than a safety afterthought. The unit of investment is not a licence. It is a course.
2. The gains concentrate in the weakest — and that is where a technical university’s losses are
Across every well-designed study in this library, the benefit of AI assistance is largest for the least expert and smallest for the most expert. For an institution losing a third of its first-year cohort, this is the most valuable finding here.
The pattern is extraordinarily consistent across domains that share nothing else. Brynjolfsson, Li and Raymond’s NBER field study of 5,179 customer-support agents found an average productivity gain of +14% issues resolved per hour — but +34% for novice and low-skilled workers, and minimal impact on experienced high-skilled workers. Their interpretation, which the data supports, is that the model diffuses the tacit knowledge of the best performers to everyone else. Stanford’s Tutor CoPilot trial found the same shape inside education: +4pp overall, but +9pp for students working with lower-rated tutors — the AI closed part of the gap between a weak tutor and a strong one. The Dell’Acqua, Mollick and Lakhani field experiment at BCG, run on 758 consultants, found the largest gains among below-average performers, who improved dramatically toward the group mean.
This is a compression effect, and it points somewhere very specific. A technical university’s most expensive, most persistent and least discussed failure is not the mediocrity of its best students. It is attrition in the first two years, concentrated in the gateway subjects — mathematical analysis, physics, mechanics, the first programming course — where a student who falls two weeks behind cannot recover because the only remediation available is a queue outside an office hour that clashes with another lecture. The published Czech numbers make the scale concrete: at ČVUT, first-year bachelor study failure ran at 31.8% university-wide in 2024, ranging from 9.5% at the Faculty of Architecture to 51.1% at the Faculty of Mechanical Engineering and 48.1% at Transportation Sciences. Half of one faculty’s incoming bachelor cohort fails in year one. Those are not students who lack capacity; they are overwhelmingly students who lacked a patient explanation at eleven at night in week six.
That is precisely the good that a constrained tutor supplies at negligible marginal cost, and it is precisely the population the evidence says benefits most. It is worth being blunt about the arithmetic, because it reframes the entire investment case. A technical university spends heavily on admissions marketing to fill a cohort, then loses a third of it to a failure mode that a well-built tutoring layer measurably addresses. The retention gain is worth more, in both money and mission, than every productivity saving in the administrative use cases that dominate university AI strategies.
Two cautions before anyone builds the business case. First, compression is not only good news: if AI lifts the floor without lifting the ceiling, the signal value of a degree compresses too, and the strongest students gain least from the institution’s investment. Second, the same compression logic that helps a struggling first-year is what makes over-reliance dangerous later — a student permanently held at the level the tool provides has been given a floor and a ceiling in the same object. Finding five deals with that directly.
What the evidence does not show. None of the compression studies are from engineering education specifically, and the two largest are from work settings rather than degree programmes. They measure task performance, not the acquisition of durable expertise — and the mechanism, diffusion of expert tacit knowledge to novices, is precisely the mechanism that could substitute for learning rather than produce it. No study in this library demonstrates that AI tutoring reduces university dropout. That trial has not been run, which is itself a finding, and it is exactly the trial a technical university with a 31.8% first-year failure rate is best placed in Europe to run.
What it obliges. Point the first serious deployment at the gateway courses with the worst failure rates, not at the flagship master’s programmes where the enthusiasts are. And instrument it as a trial, so the institution ends up owning evidence rather than anecdote.
3. Detection does not work — assessment redesign is the only durable lane
AI-text detection has been systematically tested and it fails; worse, it fails asymmetrically against non-native speakers. Any integrity policy resting on detection is both ineffective and an equity liability.
Weber-Wulff and colleagues, working through the European Network for Academic Integrity, tested fourteen AI-text detection tools under controlled conditions. The finding is unambiguous: the tools are neither accurate nor reliable, they skew toward classifying output as human-written, and their performance collapses under light obfuscation — machine translation, minor paraphrase, a pass through another model. This is not a maturity problem that a better product cycle fixes; it is close to information-theoretically inevitable as models converge on fluent, unremarkable prose.
The equity finding is worse and should end the argument on its own. Liang, Yuksekgonul, Mao, Wu and Zou at Stanford ran GPT detectors over TOEFL essays written by non-native English speakers and over essays by native-speaking US eighth-graders. The detectors were near-perfect on the native-speaker writing and misclassified roughly 61% of the non-native-speaker essays as AI-generated, with over half flagged by all seven detectors tested. The mechanism is that detectors key on lexical richness and syntactic variety — exactly the dimensions on which a competent second-language writer differs from a native one. For any technical university with a substantial international cohort, deploying such a tool means systematically accusing international students of misconduct at several times the rate of domestic ones, on the basis of their second-language fluency.
So the enforcement lane is closed. The library’s answer to what replaces it is unusually well-developed, because Australia’s regulator did the work first. TEQSA’s 2023 discussion paper Assessment Reform for the Age of Artificial Intelligence set out the two principles that now underpin most serious sector guidance worldwide, and its 2025 follow-up, Enacting Assessment Reform in a Time of AI, reports what institutions actually did with them. The core move is to stop treating every assessment as if it served one purpose. Some assessment exists to certify that a named human has a capability — and that requires secured conditions and identity assurance, at programme level rather than in every task. Everything else exists to develop capability — and there AI use should be open, taught and part of the point. Trying to make every assignment do both jobs is what produced the unwinnable arms race. QAA’s guidance adds the governance procedure: a four-step triage of which assessments actually need redesign, run through existing internal quality assurance rather than as a parallel emergency process.
For engineering, the practical translation is more favourable than for most disciplines, and the library supplies the specifics. The CDIO paper on project-based assessment in the era of generative AI reworks the PBL evaluation grid with explicit AI-use criteria — you grade the process, the design decisions and the defence, not only the artefact. The Integrevise research report on oral assessment gives the operating model, cost and staffing constraints of running vivas at cohort scale, which is the obvious authentication lane for a technical university and the one most often dismissed as impossible without checking the numbers. And ČVUT’s own Methodological Instruction 5/2023 already contains an activity-by-activity permitted / partly-permitted / forbidden schema — a more concrete instrument than most European universities possess, though written before the assessment-redesign literature matured and now due a revision that moves it from a rules document to a design document.
What the evidence does not show. It does not show that detection tools are useless in every configuration — they retain some signal on unedited long-form output, and Turnitin-class vendors dispute the specific error rates. It does not show that secured in-person assessment is unproblematic: TEQSA is explicit that identity assurance at programme level is hard, expensive and easy to implement badly. And there is no strong evidence yet on whether two-lane assessment actually preserves standards, because it is too new to have graduated a cohort.
What it obliges. Retire detection as a basis for misconduct proceedings, immediately and explicitly, and say why in public so that staff stop relying on it informally. Then run the triage: for every programme, identify the small number of points where certification genuinely requires secured conditions, secure those properly, and free everything else to be taught with AI in the open.
4. Computing education is the most disrupted subject on earth — and the disruption is to objectives, not to cheating
The subject a technical university teaches most confidently is the one large language models are best at. The published response from the computing-education research community is not about misconduct; it is about which learning objectives are still worth certifying.
No other discipline has responded to generative AI with this much empirical work this fast, which makes computing education the closest thing the sector has to a natural experiment. The ITiCSE working group report — twenty-odd authors, the landmark community survey in this library — covers code generation, code explanation, autograders, automated feedback, integrity and curriculum response in a single document, and its conclusion is that the pedagogical questions dwarf the disciplinary ones. Becker, Denny, Finnie-Ansley, Luxton-Reilly, Prather and Santos put the point in the title of their SIGCSE paper: Programming Is Hard — Or At Least It Used To Be. Their argument is that a large fraction of CS1’s traditional learning objectives were proxies. We never actually wanted students to be able to write a for-loop from memory; we wanted them to be able to decompose a problem, and writing the loop was how we checked. The proxy has broken. The underlying objective has not.
The industry-side evidence explains why the objectives must move rather than be defended. Peng, Kalliamvakou, Cihon and Demirer’s randomised controlled trial — GitHub, Microsoft and MIT — found developers using Copilot completed a standard programming task about 56% faster than the control group. Whatever a university thinks about AI in coursework, its graduates enter a profession where this is the baseline expectation. Certifying an ability to produce code unaided, slowly, is certifying a skill the employer will not buy.
The most useful study in this angle is the most uncomfortable. Prather and colleagues observed CS1 students actually using Copilot on a real assignment — the paper is titled, from a student quote, “It’s Weird That it Knows What I Want”. What they document is a set of genuinely new interaction pathologies: drift, where the student’s mental model of the problem quietly diverges from the code accumulating on screen; over-trust, where plausible output is accepted without verification; and a collapse of the metacognitive loop that novice programming is supposed to build. Ma, Chen and Konomi’s study of dialogue logs from a beginner Python course adds the typology, clustering student–ChatGPT interaction into four distinct usage patterns and linking each to performance — which means usage pattern, not usage volume, is the variable that matters and the thing worth teaching.
The instructional-response literature is thinner but concrete. The ASEE study on ChatGPT in programming courses documents the practice of requiring students to submit their own code alongside the AI’s and account for the difference — an assessment pattern that converts the tool into the object of study. The JITE paper supplies the instructor-side view of benefits and adverse impacts, which is what faculty development has to start from.
What the evidence does not show. There is still no strong longitudinal evidence on what happens to programming expertise across a whole degree under heavy AI use — the studies are single-course, single-semester, and mostly measure task outcomes rather than developed capability. The Copilot productivity RCT measured a well-specified task, not the messy comprehension-and-maintenance work that dominates real engineering. And the four-pattern typology is from one course at one university.
What it obliges. Rewrite CS1 and CS2 learning outcomes explicitly around decomposition, specification, verification, debugging and reading unfamiliar code — the objectives that survive — and assess those directly rather than through the broken proxy. Then treat every other engineering discipline as being roughly two years behind computing on the same curve, and start the same work now rather than waiting for its own crisis.
5. Over-reliance is measurable, and it is the real cost — not plagiarism
The strongest argument against casual AI adoption is not integrity. It is that confident use of a capable model measurably reduces the critical-thinking effort of the person using it — and students are more worried about this than their teachers are.
The central study is Lee and colleagues at Microsoft Research and Carnegie Mellon, published at CHI 2025: 319 knowledge workers supplied 936 first-hand examples of generative AI use at work. Two findings matter. Higher confidence in the AI predicts less critical-thinking effort; higher self-confidence in one’s own expertise predicts more. And the nature of the effort shifts — from information gathering and problem-solving toward information verification and response integration. That is not automatically bad: verification is real cognitive work and, done well, is exactly the skill finding ten of this report argues should be taught. It is bad when it is not done, and the study’s confidence finding says that the better the tool gets, the less likely the user is to do it.
This connects to a body of work that ENSI’s Education for the Agentic Age library treats at length under cognitive debt and cognitive offloading — the well-established finding that capability which is habitually externalised is not merely unused but progressively unavailable. In a professional formation context, that is the whole risk. An engineer who cannot check the model is not an engineer who is slower; they are an engineer who cannot tell when the answer is wrong, in a profession where being unable to tell is how people get hurt.
The most striking evidence that this is not an academic worry comes from students themselves. MIT’s Ad Hoc Committee report of August 2026 cites a survey of 1,002 affiliates in which 90% were somewhat or very concerned about overreliance, including 67% “very concerned”, and 45% very concerned about inaccurate outputs — this among a population where 46% of undergraduates use LLMs daily. Students are simultaneously the heaviest users and the most alarmed constituency. The institution-wide MIT survey found only 23% of respondents optimistic about generative AI, and a split verdict on self-efficacy: 40% felt AI made them more capable against 27% who felt it made them more replaceable.
The counterweight in this library is the complementarity literature, which is more sober than the enthusiasm it is usually cited to support. Hemmer and colleagues’ review formalises when a human–AI team actually beats either party alone and then reports the empirical record, which is frequently disappointing — genuine complementarity is harder to achieve than the framing suggests, and many studies find the team underperforming the better of its two members. Dell’Acqua and colleagues’ jagged-frontier experiment gives the sharpest single image: inside the frontier of tasks AI handles well, consultants produced work rated 40% higher in quality; on a task just outside it, they were 19 percentage points less likely to reach the correct answer than colleagues working without AI. The tool did not merely fail to help outside its competence. It actively degraded performance, because people could not see where the edge was.
What the evidence does not show. The Microsoft/CMU study is correlational and self-reported; it establishes an association between confidence and reduced effort, not that AI use causes cognitive decline. There is no longitudinal study in this library tracking engineering students’ capability across a degree under sustained AI use — the single most important missing evidence in the entire field. And the jagged-frontier result comes from consulting tasks, not technical ones.
What it obliges. Teach the frontier explicitly — where these systems are strong, where they fail, and how to tell which side of the line a given task sits on — as a required, assessed competence rather than an induction slide. And design at least one point in every programme where the student must demonstrate the underlying capability without assistance, not to catch anyone, but because a professional formation that never checks is not a formation.
6. The deployed teaching agent is cheap and useful — and it leaks
Two universities have run educational AI agents at real scale for years and published honest evaluations including their failure rates. The cost is trivially low, the reception is excellent, and the guardrails do not hold. All three facts must be planned for together.
Harvard’s CS50 is the largest deployment on record and the most valuable document in this library for anyone about to build something. By mid-November 2024 the CS50 Duck had been used by approximately 211,000 students, processing 10 million queries, at an average cost of $1.50 per student per year. The rollout was staged sensibly — 70 students in summer 2023, several hundred on campus that autumn plus thousands online — and the reception was strong: 75% of students used the tools frequently, 94% found them helpful and effective; in the earlier cohort, 53% “loved” and 33% “liked” them. An accuracy audit in that earlier paper found 22 of 25 curricular answers correct (88%) and 30 of 39 administrative answers correct (77%).
Then the failure data, which CS50’s team deserves considerable credit for publishing. Of the 10 million messages, roughly 2.1 million responses — 22% of all interactions — contained code blocks despite the system being instructed not to hand out solutions; at conversation level the figure is 48%, meaning about 635,000 of 1.3 million conversations involved the agent generating code. The team names the mechanism: instruction dilution, where a long conversation progressively erodes the system prompt’s authority until the pedagogical constraint stops binding. Their response was not a better prompt but a human-feedback loop with teaching fellows reviewing and correcting agent behaviour — which is to say the working architecture is not an agent, it is an agent plus a staffed quality process.
Georgia Tech’s Jill Watson provides the longer time series and the operational numbers. Across the OMSCS programme — an online master’s of roughly 9,000 students that at one point supplied about 9% of all US computer-science master’s degrees — Jill Watson Q&A served more than 4,000 students across more than a dozen classes and saved teachers more than 500 hours. Quality improved steeply with iteration: coverage rose from about 21% of questions at 80% precision in 2017 to over 96% coverage at over 86% precision by autumn 2019. Agent Smith, the companion system, generates a new Jill Watson for a fresh syllabus in about 25 hours of preparation. The LLM-era rebuild reports 76.7% answer accuracy against 31.3% for a generic OpenAI Assistants baseline on the same evaluation, with a 6.8-second average response time — the clearest available demonstration that grounding in course material, not model choice, is what makes an educational agent work.
What the evidence does not show. Neither deployment is a controlled trial: nobody randomised students into having the duck or not, so the learning effect is unmeasured and the satisfaction numbers cannot substitute for it. Both are computer-science courses at elite institutions with unusual engineering capacity. And the cost figures are inference costs — they exclude the staff time that both papers show is the actual requirement.
What it obliges. Budget for the humans. The published evidence says a course agent costs a couple of euros per student per year in compute and a permanent fraction of a teaching-assistant post in supervision — and that the supervision is what separates Georgia Tech’s 96% coverage from a broken pilot. Institutions that fund the licence and not the loop are buying the leak.
7. Faculty capability is the binding constraint — and the frameworks already exist
Every serious study of why AI adoption stalls in universities returns the same answer: not technology, not policy, not student resistance — staff capacity, incentives and time. The instruments to fix it are published, free, and largely unused.
Ithaka S+R’s interview study across nineteen North American universities is the most textured account of what actually happens: instructors adopt in isolated pockets, driven by individual enthusiasm, hitting institutional frictions that have nothing to do with the technology — no time to redesign a course, no recognition for doing so, no clarity on what is permitted, and no one to ask. The arXiv multi-institution study of barriers to generative AI adoption maps the same phenomenon across disciplines and roles, and its value is in showing the barriers are multi-level: individual confidence, departmental norms and institutional policy each block independently, so fixing one changes little. The EDUCAUSE landscape data shows institutions know this — training for faculty (63%) and staff (56%) top the list of AI strategic-planning elements — while the same sector’s Horizon Report describes faculty roles and workload as the variable most likely to determine outcomes.
What makes this finding actionable rather than merely gloomy is that the instruments exist. UNESCO’s AI Competency Framework for Teachers specifies fifteen competencies across five dimensions — human-centred mindset, ethics of AI, AI foundations and applications, AI pedagogy, and AI for professional development — at three progression levels. The European Framework for the Digital Competence of Educators (DigCompEdu), from the Commission’s Joint Research Centre, gives twenty-two competences in six areas on a six-level proficiency ladder, and is the instrument a European public university’s staff development is already legible against. And ETH Zurich’s AI Competence Framework for Lecturers is the one built by a peer technical university for its own academics — nine pages, organised by competence area and proficiency level, and effectively a ready-made curriculum that any European technical university could adapt in a term rather than draft in a year.
The gap between the availability of these frameworks and their use is the finding. Sector surveys report institutions planning training; the interview evidence reports academics who have never been offered any. A framework that is cited in a strategy document and never converted into scheduled, workload-credited, discipline-specific development produces nothing.
What the evidence does not show. There is no evaluation in this library of a faculty AI-development programme against a learning or teaching-quality outcome — the frameworks are consensus instruments, not tested interventions. The ERIC survey evidence on faculty technology use predates generative AI. And the barriers study is a preprint from a single multi-institution sample.
What it obliges. Treat academic AI capability as an operational deliverable with a named owner, a budget and a completion figure — not as a set of optional workshops. Adopt ETH Zurich’s lecturer framework rather than writing one. And make the entry point discipline-specific: a mechanical engineer will not attend a generic session on prompt writing, and should not be asked to.
8. Engineering’s artefact-based pedagogy is a structural advantage nobody is exploiting
The assessment form the rest of higher education is now scrambling to invent — work judged against something that must actually function — is engineering’s native mode. The sector’s problem is that it has been drifting away from it for thirty years.
The assessment-reform literature converges on a single prescription: make the object of assessment something a model cannot supply, which in practice means process, defence, iteration and a working result. Engineering already has all four in its lab, its design studio and its capstone. A structures calculation is checked against statics. A control loop either stabilises or oscillates. A robot completes the task or falls over. The verification is external to the assessor’s impression of the text — which is precisely the property that essay-based disciplines have lost and cannot easily rebuild.
The library’s engineering-specific material shows what exploiting this looks like. The CDIO conference paper on project-based assessment in the era of generative AI reworks the PBL evaluation grid with explicit AI-use criteria, giving a concrete rubric pattern for grading team projects where AI is permitted — the assessment is of design decisions and their justification, not of authorship. The ASEE study of image-generative AI inserted into conceptual design in a CAD class is the closest thing to a design-studio protocol for AI-assisted ideation. Martín-Núñez and Díaz Lantada’s review in the International Journal of Engineering Education supplies the taxonomy that keeps the strategy coherent: AI as subject matter to be taught versus AI as teaching infrastructure to be built — two different programmes, routinely conflated, requiring different owners and different budgets. The gAI-PT4I4 paper on generative AI plus low-fidelity digital twins with VR and retrieval-augmented generation is the strongest available template for scaling laboratory and simulation work when physical lab capacity is the constraint.
The peer-institution consensus documents matter here because they establish that this is the declared direction of European technical universities rather than one institution’s bet. CESAER’s Engineer of the Future white paper, written by the association of European universities of science and technology, argues for competence-based and challenge-based learning, lifelong learning and digital transformation as the shape of engineering education. EuroTeQ’s Framework of Qualifications — the alliance deliverable defining what a European engineering graduate must be able to do — is the concrete competence architecture into which AI competences can be inserted without inventing a parallel framework. Both are network commitments a member institution has already made and can simply execute against.
The uncomfortable half of this finding is that the advantage is being squandered. Artefact-based assessment is expensive in staff time, and three decades of expanding cohorts with flat teaching budgets have pushed engineering programmes steadily toward the cheap end — the individually-submitted problem set, the templated lab report, the multiple-choice test. Every one of those is now worthless as evidence of capability. The strategic point is that the response to generative AI and the response to the long erosion of engineering pedagogy are the same response, which makes the current moment an unusually good one to argue for the resources.
What the evidence does not show. There is no controlled evidence that project-based assessment resists AI better than written assessment — it is an argument from the nature of the task, not a measured result, and a team project can be substantially AI-generated in ways a busy assessor will not detect. The CDIO and ASEE studies are small, single-course and self-reported. And the digital-twin work is a proposed framework with limited deployment evidence.
What it obliges. Audit where each programme’s assessment actually sits on the spectrum from artefact-and-defence to submitted text, and move the balance deliberately — accepting that this costs contact hours and saying so in the budget rather than pretending redesign is free.
9. The law has already decided most of this is high-risk
The EU AI Act classifies as high-risk exactly the university uses an efficiency-minded administration would automate first. This is not a future compliance question; the obligations are in force and the sequencing consequence is immediate.
Annex III of Regulation (EU) 2024/1689 names education and vocational training explicitly. The high-risk categories cover AI systems used to determine access or admission to educational institutions, to evaluate learning outcomes, to assess the appropriate level of education a person will receive, and to monitor and detect prohibited behaviour during tests. Read that list against a typical university AI wish-list — automated admissions triage, automated grading, adaptive placement, remote proctoring — and the overlap is nearly total. High-risk classification brings risk management, data governance, technical documentation, logging, transparency, human oversight and accuracy/robustness obligations, and a provider-versus-deployer distinction that changes materially if the university builds rather than buys.
This is why the sequencing in the earlier findings is not merely pedagogically sound but legally forced. Tutoring, explanation, feedback-for-learning and staff support are not on the Annex III list. Grading, admission and proctoring are. An institution that starts with the low-risk teaching layer builds capability, evidence and trust while its compliance work matures; an institution that starts with automated grading because it looks like the biggest efficiency win has taken on the heaviest obligations first, with no institutional competence yet built.
Three further instruments define the operating envelope, and they are cumulative rather than alternative. The ESG — the Standards and Guidelines for Quality Assurance in the European Higher Education Area — govern how any change to teaching, assessment or grading must pass through internal quality assurance and survive external accreditation; an AI-mediated assessment change that has not been through programme-level QA is not merely risky, it is unaccredited. Data protection sits on top: the EDPS orientations on generative AI and the EUDPR, and the GDPR regime generally, apply in full to student data flowing through a commercial model. Jisc’s Code of Practice for Learning Analytics is the practical instrument that converts those obligations into institutional procedure — responsibility, transparency, consent, minimising adverse impact and stewardship — and remains the best available starting point for a university’s data-governance model.
The Czech and European policy frame supplies the funding logic rather than further constraint. The Czech National AI Strategy to 2030 carries education, skills and research pillars and is the national mandate a public university’s programme can be attached to; the Government Council’s 2024 analysis documents that the country has no dedicated digital-education strategy, which is both a gap and an opening for an institution willing to define the practice. The Commission’s Digital Education Action Plan and the joint OECD/EC AI literacy framework provide the vocabulary against which national and EU funding will judge proposals.
What the evidence does not show. The AI Act’s application to specific university configurations is genuinely unsettled — whether a university fine-tuning a commercial model becomes a provider, how the research exemption interacts with teaching deployments, and where the boundary sits between a feedback tool and an outcome-evaluation system are all live questions that guidance has not fully resolved.
What it obliges. Sequence deliberately by risk class, not by perceived efficiency. Put the AI Act’s high-risk register, the ESG and the data-protection regime into one governance instrument owned by one office, before the first pilot rather than after the third.
10. Almost none of this has been evaluated to the standard the university demands of its own research
A technical university applies rigorous evidence standards to everything except its own teaching. The AI literature is where that inconsistency becomes expensive.
The standards exist and are in this library. The What Works Clearinghouse Procedures and Standards Handbook sets out how the US government decides whether an education study counts as evidence — design requirements, attrition thresholds, baseline equivalence, effect-size computation. The EEF’s evaluator guide is the operational playbook: protocols, pre-registration, statistical analysis plans, implementation and process evaluation. Measured against either, the great majority of published claims about AI in higher education — including a large share of what circulates as best practice — would not qualify as evidence at all. They are satisfaction surveys, single-course reflections, and vendor case studies with no comparison condition.
The specific technical caution matters most for the intervention universities most want to build. Gardner and colleagues’ study of temporal and between-group variability in college dropout prediction shows early-warning model performance degrading across cohorts and across student subgroups. A model trained on last year’s students underperforms on this year’s, and underperforms unevenly — worse for the subgroups an equity-minded institution most wants it to serve. Since at-risk prediction is both the most attractive analytics application and an Annex III-adjacent use, this is the finding that should govern how it is deployed: as a trigger for offering help, monitored for subgroup drift, never as an input to a decision about the student.
The rest of the field’s methodological weaknesses are ordinary and predictable. The meta-analytic g = 0.670 pools mostly short interventions with proximal outcomes, exactly the conditions under which effect sizes shrink on replication. Deployment papers report satisfaction and usage, not learning. Almost nothing measures retention beyond the end of the intervention. And the studies with the best designs — Harvard’s physics trial, the Nigeria RCT, Tutor CoPilot — are, tellingly, the ones reporting the most disciplined interventions, which raises the possibility that design quality and intervention quality are correlated and the field’s average effect is inflated by neither being present.
The opportunity in this is larger than the caution. The gap between what is claimed and what is established is wide, the population needed to close it sits in every gateway lecture theatre, and a technical university already employs the statisticians. An institution that instruments its own deployment as a trial ends up owning evidence that nobody else has — publishable, fundable, and a durable reputational asset in a field where almost everyone is guessing.
What it obliges. Treat every AI teaching deployment as a study: comparison condition, pre-registered outcome, learning measure rather than satisfaction measure, subgroup analysis by default. It costs little more than doing it badly and produces something the institution can defend.
What to do first
The ten findings collapse into a short sequence, and its order is not arbitrary — it runs from highest evidence and lowest regulatory risk to the reverse.
First, close the detection question and open the assessment question. Withdraw detection from misconduct procedure explicitly and publicly, then run the TEQSA/QAA triage across every programme to identify the minimum set of points where certification genuinely requires secured conditions. This costs nothing, ends an unwinnable and inequitable enforcement effort, and is the precondition for everything else.
Second, build one constrained tutor for one gateway course with a bad failure rate — grounded in the actual course material, engineered to withhold, supervised by a named teaching assistant, instrumented as a randomised trial with a learning outcome. The evidence says the effect will be largest exactly there. Budget two euros per student per year for inference and a fraction of a post for the loop, because the published deployments say the loop is what works.
Third, fund academic capability as an operational deliverable. Adopt ETH Zurich’s lecturer competence framework rather than drafting one, map it onto DigCompEdu for European legibility, deliver it discipline-by-discipline with workload credit, and publish the completion figure. Every study of stalled adoption points here.
Fourth, move the assessment centre of gravity back toward the artefact and its defence — the direction CESAER and EuroTeQ have already committed European technical universities to, now with a second and more urgent justification.
Fifth, put the AI Act, the ESG and data protection into a single governance instrument before the first high-risk pilot, and sequence deployments by risk class rather than by perceived efficiency.
The reframe to hold, against a sector conversation that will keep pulling toward tools and rules: the capability at stake is the institution’s. The models are a commodity available to every student for the price of a coffee, and no policy will change that. What is not a commodity is a university that knows which of its assessments still mean anything, whose academics can teach with and against these systems, that catches its struggling first-years before they leave, and that can prove any of it. That is buildable, the evidence says roughly how, and — the one genuinely urgent point in this report — the institutions that start now will be the ones producing the evidence everyone else cites in three years.



