Values are trained the way strength is trained — by reps under load with a coach watching — and this is the operating manual: a 32-value canon, a six-step formation loop, eight program areas, the KPIs, and the first twelve months.
Written by the ENSI Foresight Division. Built on a library of 150 primary documents downloaded from the world’s leading institutions. Compiled 2026-07-25.
The argument: formation is a capability, not a course
Nobody has ever become strong by attending lectures on strength. The gymnasium does not explain muscles to you; it puts you under a load slightly heavier than you can comfortably carry, watches your form, corrects it, and makes you do it again — and again, on a schedule, for years. The muscle is not informed into existence. It is provoked into existence by stress, feedback and repetition. Every serious tradition of human formation has known that character works the same way. Aristotle said it first and said it plainly: we become just by doing just acts, brave by doing brave acts — virtue is acquired by practice and habituation, not instruction alone (Aristotle, Ross trans. 1925; angle 02). Modern philosophy has caught back up: the strongest current account treats virtue as a skill, acquired the way expertise is acquired — through the novice-to-expert progression of attempt, error, correction and refinement (Stichter 2007; angle 02), a framing that moral psychology now shares in Narvaez’s model of ethical expertise development (Narvaez 2006; angle 03).
And yet almost every state that says it wants citizens of character runs values education as a lecture hall. A subject on the timetable. A poster in the corridor. An assembly about kindness. The library beneath this report says, from fifteen separate angles, that this cannot work — because moral judgment is intuition-first and reasoning-second, so exhortation reaches the part of the mind that does not drive behaviour (Haidt 2001; angle 03); because behaviour is carried by habits formed through repetition in stable contexts, not by conclusions (Wood 2016; angle 07); and because the most rigorous multi-programme evaluation ever run on packaged school character curricula — bolt-on lessons added to an unchanged school — found essentially nothing (US IES 2010; angle 15). The lecture model has been tested at scale. It failed.
Report 1 of this series established the evidence. This report is the operating manual. It answers the “so what, now what” question for the actor that matters: the state and its institutions — ministries, schools, academies, professional bodies. The state is the right actor for the same reason it is the right actor for public health: the returns are enormous, long-horizon, and diffuse. Childhood emotional and character formation predicts adult life satisfaction better than academic achievement does (LSE CEP 2013; angle 10); non-cognitive skills carry causal weight for employment, health and civic life that credentials do not capture (Heckman & Kautz 2014; angle 13). Singapore already runs the world’s most systematized national character curriculum (Singapore MOE 2021; angle 15); Japan has taught dōtoku for generations (NIER 2013; angle 15). The question facing a European state is not whether its institutions form character — they do, every day, by accident, through what they reward and tolerate — but whether they will do it deliberately, transparently and well.
Why now — three reasons. First, the evidence is mature. We no longer have only the aspiration; we have the mechanism. A meta-analysis of 213 school programmes shows social-emotional formation moves both behaviour and an 11-percentile achievement gain (Durlak 2011; angle 09). Productive failure shows why struggling before instruction deepens learning (Kapur 2015; angle 05). Habit science shows how practice becomes permanent disposition (Wood 2016; Gardner 2012; angle 07). Mentoring has randomized-trial evidence (PPV 1995; angle 11). The parts exist; nobody has assembled the machine. Second, AI makes situational practice scalable. The eternal bottleneck of experiential formation was adult attention — one coach can watch only so many reps. Generative agents can now populate believable practice situations (Park et al. 2023; angle 14), LLM role-play systems let a learner rehearse a hard conversation and get feedback (Shaikh et al. 2023; angle 14), and a human-AI tutoring copilot has already scaled expert pedagogical moves across 900 tutors in a randomized trial (Wang et al. 2024; angle 14). Report 3 details this engine; this report assumes it. Third, the purpose crisis. Purpose — a stable intention to accomplish something meaningful to the self and consequential beyond it (Damon 2003; angle 12) — is measurably linked to health, achievement and persistence (Templeton/Bronk 2020; angle 12), and benevolence measurably moves national happiness (WHR 2025; angle 10). Societies with meaning deficits do not lack information. They lack formation.
So this report builds the gymnasium. It names a canon of 32 values in eight families — explicitly swap-ready, because the naming is the institution’s sovereign act. It specifies the six-step formation loop that constitutes a single rep. It ranks eight programme areas that together are the training floor. It sets out measurement that will not be destroyed by its own stakes, governance that keeps formation from sliding into indoctrination, and a first twelve months for a mid-sized EU state — the Czech Republic is our worked example throughout. The thesis in one line: formation is a capability an institution builds and runs, not a course it schedules — and the states that build it first will compound trust, honesty and competence the way early adopters of public schooling compounded literacy.
The playbook in brief
Character is a trainable skill, not transferable information — virtue is acquired the way expertise is acquired: attempt, error, correction, repetition (Aristotle trans. 1925; Stichter 2007; angle 02; Narvaez 2006; angle 03).
The canon comes first. Thirty-two values in eight families — Truth, Courage, Discipline, Justice, Wisdom, Love, Strength, Purpose — each defined as an observable trained behaviour with a signature training situation. The list is swap-ready; the discipline of naming is not.
One rep = the six-step formation loop: Situate → Attempt → Fail productively → Correct → Habituate → Prove. Each step carries its own evidence base, from VR dilemmas (Francis et al. 2016; angle 06) to productive failure (Kapur 2015; angle 05) to implementation intentions (Gollwitzer & Sheeran 2006; angle 07) to situational judgment tests (Europe PMC/PLOS 2019; angle 13).
Eight programme areas make the gymnasium, ranked by evidence and leverage: whole-school ethos, the challenge curriculum, the dilemma gym, the failure ladder, the habit protocol, mentorship infrastructure, purpose-matching, and adult/professional formation.
Culture beats curriculum. Bolt-on character programmes returned null results in the largest randomized evaluation (US IES 2010; angle 15); embedded whole-school strategies with SAFE design features are what works (Durlak 2011; SRCD 2012; angle 09; Berkowitz, Bier & McCauley 2017; angle 01).
Measurement must be multi-method and never high-stakes — self-report is corroded by reference bias and faking (RAND 2014; Heckman & Kautz 2014; angle 13); combine situational judgment tests, observed behaviour, 360-degree views and longitudinal outcomes; measure growth, not rank; the learner owns the data.
Governance is the licence to operate: a transparent, contestable canon; judgment trained, not compliance; the Japanese and Singaporean state-curricula lessons heeded (NIER 2013; Singapore MOE 2021; angle 15); Kristjánsson’s answers to the standard objections on the record (Kristjánsson 2013; angle 02).
The failure mode is institutional hypocrisy. An institution that demands the impossible teaches its members to lie — the US Army documented this on itself (Wong & Gerras 2015; angle 08). The gymnasium must audit its own demands.
Twelve months to first proof for a state the size of the Czech Republic: name the canon, pick 20 pilot schools and 2 professional academies, train mentors to published standards, stand up a dilemma-gym MVP, baseline with serious instruments, publish everything openly.
The canon — 32 values in eight families
Every gymnasium trains named lifts. A formation system trains named values — and the naming is the first act of seriousness, because an unnamed value cannot be practised, coached or measured. What follows is ENSI’s proposed canon: 32 values in eight families. It is deliberately swap-ready — a ministry, an academy or a school should fight about this list, strike values, add its own, and then commit in public. The framework survives any reasonable substitution; what does not survive is vagueness. Each value below is given as an operational definition — what the trained behaviour looks like, since a value that cannot be seen cannot be trained — and the signature training situation in which it is most directly exercised. The situations are drawn from the methods evidenced across this library; the loop in the next section shows how each becomes a rep.
Family 1 — Truth. The load-bearing family: without honest reporting there is no feedback, and without feedback nothing else in this playbook works.
Honesty — says what is true when a lie would be cheaper. Signature situation: the after-action review in which the learner names their own error before anyone else can.
Integrity — behaviour matches stated values when nobody is watching. Signature situation: unproctored work under a real honor code with real stakes (ERIC 2010; angle 08).
Sincerity — means what it says; refuses the performance of virtue. Signature situation: the feedback circle where flattery is challenged and plain speech is rewarded.
Accountability — owns outcomes, including failures, without excuse or deflection. Signature situation: carrying real responsibility for a project whose results are publicly reviewed.
Family 2 — Courage. The family that converts conviction into action under fear or uncertainty.
Boldness — acts decisively under uncertainty instead of waiting for permission. Signature situation: the time-pressured simulation or expedition decision that cannot be deferred (Sütfeld et al. 2017; angle 06; AIR 2005; angle 04).
Moral courage — names a wrong at social cost. Signature situation: the “Should I Say Something?” rehearsal — confronting a lapse by a peer or a superior in simulation before doing it in life (Europe PMC 2023; angle 06).
Initiative — starts without being told. Signature situation: the open-brief project where the problem itself must first be found (MDRC 2017; angle 04).
Enterprise — builds something new that others actually use. Signature situation: running a real venture or service with real users and real consequences.
Family 3 — Discipline. The family that makes every other value repeatable on a bad day.
Self-control — chooses the harder-better over the easier-worse in the moment. Signature situation: the daily WOOP rep — wish, outcome, obstacle, plan (Duckworth et al. 2011; angle 07).
Diligence — sustains careful effort past the point of boredom. Signature situation: logged deliberate-practice blocks with a coach watching form (Ericsson et al. 1993; angle 07).
Order — keeps spaces, systems and commitments structured. Signature situation: maintained personal standards inspected without warning — the cadet-room discipline (West Point 2019; angle 08).
Patience — tolerates delay without abandoning the goal. Signature situation: the long-horizon project whose results only appear after months.
Family 4 — Justice. The family that orients power and advantage toward the right.
Righteousness — orients to what is right over what is advantageous. Signature situation: the structured dilemma discussion where the right and the profitable collide (Lind 2021; angle 06).
Fairness — allocates by consistent principle, not by favour. Signature situation: the resource-allocation role-play in which the learner’s own side loses under the fair rule.
Respect — treats every person as an end, especially the inconvenient ones. Signature situation: the restorative circle and the structured encounter across social difference.
Civic duty — contributes to the commons unprompted. Signature situation: service-learning with structured reflection and real community stakes (Celio et al. 2011; angle 04).
Family 5 — Wisdom. The family Aristotle put in charge: phronesis, the executive that decides which value the situation demands (Jubilee Centre 2022; angle 01).
Practical wisdom — reads the situation and chooses the right act among competing goods. Signature situation: adjudicating live cases with a mentor, exercising all four components — perception, adjudication, emotion, identity (Jubilee Centre 2020; angle 02).
Curiosity — asks the next question unbidden. Signature situation: open inquiry where the syllabus does not contain the answer.
Discernment — separates signal from noise and truth from spin. Signature situation: the exercise seeded with misleading information — the wargame with a lying source (Emery/TNSR 2021; angle 06).
Foresight — acts today against consequences years out. Signature situation: the scenario exercise in which decisions are replayed against unfolding futures.
Family 6 — Love. The family that turns formation outward; benevolence is measurable and it moves whole societies (WHR 2025; angle 10).
Compassion — moves toward suffering rather than away from it. Signature situation: the sustained care placement with real dependents, not a visit.
Generosity — gives time and resources at genuine cost. Signature situation: the giving project — designed, budgeted and delivered by the learner (Sparks et al. 2019; angle 11).
Loyalty — stays with people through cost. Signature situation: the team expedition where quitting harms teammates, not just oneself.
Forgiveness — releases a genuine grievance without denying the harm. Signature situation: the restorative-justice conference, face to face with real harm.
Family 7 — Strength. The family built almost entirely on the failure ladder — these values cannot be trained without adversity.
Perseverance — continues past repeated failure. Signature situation: problems calibrated just beyond current ability, attempted before instruction (Kapur & Roll 2018; angle 05; Eskreis-Winkler et al. 2014; angle 07).
Resilience — recovers form after a real blow. Signature situation: supported adversity — the expedition that goes wrong by design, inside a safety envelope.
Humility — updates when proven wrong. Signature situation: the post-error debrief in which the learner’s confident model publicly fails (Metcalfe 2017; angle 05).
Gratitude — registers what is given rather than what is owed. Signature situation: the daily noticing practice — three good things, done until automatic (Seligman et al. 2005; angle 10).
Family 8 — Purpose. The family that answers why train at all — and the strongest known motivational engine for the rest.
Calling — connects daily effort to a contribution beyond the self. Signature situation: the purpose interview and the self-transcendent reframing of ordinary work (Damon 2003; Yeager et al. 2014; angle 12).
Hope — builds concrete pathways to a wished-for future. Signature situation: mental contrasting with implementation intentions — the wish met honestly by the obstacle and the plan (Gollwitzer & Sheeran 2006; angle 07).
Reverence — stands rightly before what is larger than the self. Signature situation: the wilderness solo; the encounter with things too large to master.
Stewardship — leaves what it touches better, and hands it on. Signature situation: owning a real asset — a garden, an archive, a younger cohort — across a full year.
Eight families, thirty-two values, every one of them stated as a behaviour and paired with a situation. That is the syllabus of the gymnasium. What follows is the rep.
The formation loop — what one rep looks like
A gymnasium is not defined by its equipment but by its unit of work: the rep, performed with correct form, under progressive load. The formation loop is the rep of character training. It compresses what this library establishes from separate directions — Kolb’s experiential learning cycle (Kolb, Boyatzis & Mainemelis 2001; angle 04), the productive-failure paradigm (Kapur 2015; angle 05), the science of habit (Wood 2016; angle 07) and the Jubilee Centre’s caught–taught–sought account of how character actually enters a person (Jubilee Centre 2022; angle 11) — into six steps: Situate, Attempt, Fail productively, Correct, Habituate, Prove. The loop is scale-invariant. It can run in a ten-minute classroom drill, a semester-long service project, or a 47-month academy programme (West Point 2025; angle 08). What may never be skipped is a step — and the step institutions always skip is the third.
Step 1 — Situate: put the learner in a situation, not in front of one
Formation begins by placing the learner inside a situation that demands the value — real where possible, simulated where reality is too dangerous or too rare. The evidence for insisting on situations rather than descriptions of situations is now direct. When moral dilemmas are presented in immersive VR rather than as text vignettes, people respond differently — simulated moral action diverges from armchair moral judgment (Francis et al. 2016; angle 06), and a growing review literature maps how VR dilemmas expose the gap between what people say and what they do (Europe PMC 2025; angle 06). Under time pressure in VR road-traffic dilemmas, ethical decision behaviour can be modelled as it actually operates — fast, embodied, constrained (Sütfeld et al. 2017; angle 06). The military reached the same conclusion in the field: ethics scenarios woven into high-intensity field exercises train what classroom ethics cannot (Europe PMC 2014; angle 06). And where full simulation is impractical, the Konstanz Method of Dilemma Discussion gives teachers an operational, manualised way to stage a genuine moral situation inside an ordinary classroom (Lind 2021; angle 06). The deeper reason situating matters is Haidt’s: moral judgment runs intuition-first (Haidt 2001; angle 03) — and intuitions are trained by encounters, not by arguments. Design rule: every value in the canon has its signature situation; the institution’s job is to manufacture encounters with it on a schedule.
Step 2 — Attempt: act under uncertainty, before being shown how
Inside the situation, the learner must act — decide, commit, produce — before anyone demonstrates the correct answer. This is where the experiential-learning evidence carries the load. Dewey’s foundational claim that education happens through experience and its consequences (Dewey 1938; angle 04) is now backed by convergent quantitative results: a 62-study meta-analysis of service-learning shows gains in attitudes toward self, school engagement, civic engagement, social skills and academic performance (Celio et al. 2011; angle 04); an 11-study meta-analysis confirms service-learning reliably increases student learning (Warren 2012; angle 04); the definitive literature review of project-based learning sets out the design principles under which acting on authentic problems works K-12 (MDRC/Condliffe 2017; angle 04); four rigorous studies show PBL raising achievement across grades and subjects, including for low-income students (Lucas Education Research 2021; angle 04); and a 66-study meta-analysis quantifies PBL’s effects on thinking skills, attitudes and achievement (Zhang & Ma 2023; angle 04). Crucially, Celio’s meta identifies reflection and student voice as moderators — attempts formed by the learner’s own choices, later reflected on, are what move character, not supervised errands. Design rule: the attempt must be genuinely the learner’s — their plan, their call, their name on it.
Step 3 — Fail productively: design the failure in, before the instruction
This is the step that separates a gymnasium from a lecture hall, and the one conventional schooling engineers out. The productive-failure programme shows that having learners generate solutions to problems beyond their current ability — and fail — before receiving instruction produces deeper learning than instruction-first sequences, through three mechanisms: activation of prior knowledge, awareness of knowledge gaps, and better encoding of the canonical solution when it finally arrives (Kapur 2015; Kapur & Roll 2018; angle 05). The effect holds even at micro-scale for procedural knowledge (Ziegler, Trninic & Kapur 2021; angle 05) and has been extended as a design paradigm to programming education (arXiv 2024; angle 05). The Bjorks’ desirable-difficulties framework generalises the principle: conditions that make practice harder and slower — spacing, interleaving, generation — improve retention and transfer (Bjork & Bjork 2011; angle 05). Errorful learning followed by corrective feedback beats errorless learning (Metcalfe 2017; angle 05). And in adult training, a 24-study meta-analysis shows error-management training — explicitly inviting errors and framing them as informative — outperforms error-avoidant training, especially for adaptive transfer to novel problems (Keith & Frese 2008; angle 05), with modern professional evidence extending it to high-stakes clinical skill (Europe PMC 2024; angle 05). The frame matters as much as the failure: a short growth-mindset intervention reframing failure as information raised achievement in a national randomized experiment of over 12,000 students (Yeager et al. 2019; angle 05). Design rule: failure must be survivable, calibrated and expected — a designed feature the learner knows is coming, not an ambush.
Step 4 — Correct: feedback, reflection, and the mentor’s debrief
An error that is never examined trains nothing — or worse, trains the error. The corrective step is where experience becomes learning. The cognitive machinery exists and is measurable: after committing errors, people spontaneously slow down and adjust — post-error slowing and accuracy adjustment are the substrate of situational self-correction (Danielmeier & Ullsperger 2011; angle 05) — and corrective feedback after error commission is precisely the condition under which errorful learning outperforms errorless (Metcalfe 2017; angle 05). Kolb’s cycle formalises the pedagogy: concrete experience must pass through reflective observation and abstract conceptualisation before it re-enters action (Kolb, Boyatzis & Mainemelis 2001; angle 04) — which is why reflection quality moderates service-learning outcomes (Celio et al. 2011; angle 04). The human form of this step is the debrief, and its canonical model is cognitive apprenticeship: the mentor models, coaches, scaffolds — and then deliberately fades (Collins, Brown & Newman 1987; angle 11). In moral formation specifically, structured dilemma discussion measurably shifts justice reasoning on the Defining Issues Test (ERIC 2015; angle 06), the instrument tradition behind forty years of documented moral-judgment growth (Thoma 2014; angle 03). Design rule: every designed failure has a scheduled debrief with a named coach; no rep ends at the error.
Step 5 — Habituate: reps until the value runs without willpower
A corrected behaviour is still a fragile behaviour. The fifth step turns it automatic, because character that depends on daily willpower is not yet character. Habit science supplies the mechanics: behaviour repeated in stable contexts, cued and rewarded, transfers control from deliberate intention to automatic response (Wood 2016; angle 07), and the practical protocol — anchor the new behaviour to an existing routine and repeat, with automaticity plateauing over roughly 66 days in the underlying research — is documented and usable (Gardner, Lally & Wardle 2012; angle 07). Implementation intentions — if-then plans binding situations to responses — convert intention into action with a meta-analytic effect of d = .65 across 94 studies (Gollwitzer & Sheeran 2006; angle 07). Combined with mental contrasting as WOOP/MCII, the technique raised adolescents’ self-disciplined studying by roughly 60% in a school experiment (Duckworth et al. 2011; angle 07) and improved grades, attendance and conduct in replication (Duckworth et al. 2013; angle 07). For the skill-like face of virtue, deliberate practice supplies the structure — effortful, feedback-rich, structured repetition (Ericsson, Krampe & Tesch-Römer 1993; angle 07) — with two honest caveats: Ericsson’s own strictures on what actually counts as deliberate practice (Ericsson & Harwell 2019; angle 07), and Macnamara’s meta-analysis showing practice explains a real but bounded share of performance variance across domains (Macnamara, Hambrick & Oswald 2014; angle 07) — so the gymnasium must engineer context, feedback and opportunity, not just hours. Schools have an evidence-backed frame for teaching the self-regulation this step requires (EEF 2018; angle 07), and the COM-B model gives designers the full checklist — capability, opportunity, motivation (Michie, van Stralen & West 2011; angle 07). Habituation, note, is not childhood-only: the Aristotelian case that it runs lifelong is explicit (Sanderse 2018; angle 02). Design rule: every value in a learner’s current formation plan has a daily or weekly rep with a context cue — and the rep survives the school holidays.
Step 6 — Prove: demonstrate it in the world, and measure it honestly
The loop closes only when the value shows up outside the gymnasium — under observation, in situations that were not staged for the learner’s benefit. Proof has two halves. The first is demonstration: real responsibility discharged in the real world — the service project delivered (Celio et al. 2011; angle 04), the standard upheld under an honor code across years (ERIC 2010; angle 08), the pattern West Point institutionalises by making cadets live and lead honourably across a 47-month experience rather than pass a course (West Point 2025; angle 08). The second is measurement that resists self-flattery: situational judgment tests that present realistic scenarios and score the chosen response, with validated instruments now existing for character-adjacent constructs like dependability (Europe PMC/PLOS 2019; angle 13); moral-dilemma instruments fielded at scale on 10,000+ UK students (Jubilee Centre 2015; angle 13); and behavioural evidence prioritised over self-report, because self-reported character is confounded by reference bias and faking (Heckman & Kautz 2014; RAND 2014; angle 13). Proof feeds the next loop: it sets the next load. Design rule: proof events are scheduled, observed, and recorded as growth against the learner’s own baseline — the full measurement architecture, and its guardrails, are set out later in this report.
Run those six steps, on a schedule, against the 32 values, for years — that is the whole method. The remainder of this playbook is the institutional machine that makes the loop run: eight programme areas, ranked.
The eight programme areas — how the ranking works
The gymnasium is built from eight programme areas. They are ranked, not listed: by the strength of the evidence behind them, by leverage — how much of the 32-value canon each one trains — and by dependency, because some areas are the load-bearing walls the others hang from. Ethos comes first because every other area runs inside it and is silently cancelled by a culture that contradicts it. The challenge curriculum and the dilemma gym are the twin training floors — real situations and simulated ones. The failure ladder and the habit protocol are the two halves of the loop institutions most reliably omit. Mentorship is the human transmission channel, purpose-matching is the personalisation layer, and adult formation extends the whole system past age eighteen, where most states currently stop. Each area is presented as a compact operating brief: In short · Why it ranks here · Methods that fit · Signals & KPIs · Institutional wiring & first moves.
Area 1 — The whole-school ethos: caught before taught
In short. The culture of the institution is the primary curriculum. Learners absorb the values the institution practises — in its corridors, staff rooms, discipline policies and small daily transactions — long before and long after any lesson about values. The first programme area is therefore not a programme at all: it is the deliberate engineering of ethos.
Why it ranks here. The evidence is unusually consistent, from both directions. Positively: the Jubilee Centre’s field-defining framework holds that character is caught through ethos and role-modelling, taught through instruction, and sought through practice — in that order (Jubilee Centre 2022; angles 01, 11). The SRCD’s landmark analysis argues social-emotional formation works as continuous whole-school strategies embedded in daily practice, not as packaged curricula (SRCD/Jones & Bouffard 2012; angle 09). Berkowitz’s research programme distils what works into design principles — the PRIMED framework — moving the field from advocacy to implementation science (Berkowitz, Bier & McCauley 2017; angle 01; NASEM 2016; angle 01), and the landmark 213-programme meta-analysis found that programmes work when they carry the SAFE features — sequenced, active, focused, explicit (Durlak et al. 2011; angle 09). Negatively: the largest randomized multi-programme evaluation of schoolwide character curricula found essentially null effects — bolt-on programmes dropped into unchanged school cultures do not move children (US IES 2010; angle 15), and KIPP’s rigorous evaluation found strong achievement effects but flat character-survey effects even in an intensely character-branded network (Mathematica 2015; angle 09). Culture is not one factor among many. It is the medium.
Methods that fit. The Eleven Principles of Effective Character Education as a whole-school self-assessment and rubric (Character.org 2014; Lickona, Schaps & Lewis 2007; angle 01) · PRIMED design principles for implementation (Berkowitz, Bier & McCauley 2017; angle 01) · the EEF’s evidence-based recommendations for embedding social-emotional learning in whole-school routine (EEF 2019; angle 09) · staff modelling — with Carr’s philosophical caution that role-modelling is a subtle mechanism that must be honest to work, not a poster campaign (Carr 2023; angle 02) · the UK’s national framework as the state-level articulation of expected provision (DfE 2019; angle 01), grounded in the empirical map of what schools actually do (NatCen 2017; angle 01).
Signals & KPIs. School-climate survey trends year on year · virtue-literacy growth, measurable at scale — the story-based Knightly Virtues programme produced substantial virtue-literacy gains across 20,000+ pupils (Jubilee Centre 2014; angle 09) · a staff-behaviour audit: does what we reward, tolerate and punish match the canon · discipline data reframed — restorative resolutions versus exclusions · an annual “ethos contradiction” register: every routine found teaching against the canon, and its fix.
Institutional wiring & first moves. Ownership sits with the head, not a coordinator — ethos cannot be delegated sideways. First moves: run the Eleven Principles self-assessment as a baseline; audit every routine with one question — what value does this actually train? — starting with detention, marking and staff meetings; write staff recruitment and development criteria that name character explicitly; kill any planned purchase of a packaged programme that arrives without a culture plan attached.
Area 2 — The challenge curriculum: real problems, real service, real weather
In short. A standing curriculum of authentic challenges — project-based learning on real problems, service-learning in the real community, and outdoor expeditions in real weather — is the gymnasium’s main training floor for Courage, Justice, Love and Strength. It is where attempts happen.
Why it ranks here. This is the best-evidenced experiential area in the library. Service-learning: a 62-study meta-analysis shows significant gains in attitudes toward self, attitudes toward school, civic engagement, social skills and academic achievement — with reflection and student voice moderating the size of the effect (Celio et al. 2011; angle 04), corroborated by an 11-study meta on learning outcomes (Warren 2012; angle 04). Project-based learning: the definitive design-principles review (MDRC/Condliffe 2017; angle 04), four rigorous studies showing achievement gains across grades and subjects including for low-income students (Lucas Education Research 2021; angle 04), and a 66-study meta-analysis quantifying effects on thinking skills, attitudes and achievement (Zhang & Ma 2023; angle 04). Outdoors: quasi-experimental evidence that residential outdoor programmes build cooperation, conflict resolution and self-esteem alongside science learning (AIR 2005; angle 04), and a systematic review of regular curricular outdoor classes showing social, academic, physical and psychological effects (Becker et al. 2017; angle 04). No other area trains as many canon values per hour.
Methods that fit. Open-brief PBL where learners find the problem before solving it (MDRC/Condliffe 2017; angle 04) · service-learning designed around Celio’s moderators — learner voice, community need, structured reflection (Celio et al. 2011; angle 04) · residential expeditions and regular outdoor classes as the delivery vehicle for Strength-family values (AIR 2005; Becker et al. 2017; angle 04) · Dewey’s criteria of continuity and interaction as the quality test for any proposed experience (Dewey 1938; angle 04).
Signals & KPIs. Share of curriculum hours that are experiential, per term · number of projects with a real external user or client · service hours completed with structured reflection attached — hours without reflection do not count · outdoor days per learner per year · project outcomes independently reviewed by the external user.
Institutional wiring & first moves. This area lives or dies on the timetable and the assessment regime. First moves: reserve a protected weekly block for challenge work; sign standing agreements with municipalities, NGOs and firms to supply real problems and real service placements; retrain assessment to accept project evidence; put every learner outdoors on expedition at least once a year, and make the expedition’s character content explicit rather than incidental.
Area 3 — The dilemma gym: simulation, role-play, VR and wargames
In short. A standing facility — part method, part scenario library, part technology — where learners rehearse morally hard situations at high frequency and zero real-world cost: dilemma discussions, role-plays, simulations, VR scenarios and wargames keyed to the 32 values. Reality is the best trainer but a rare and expensive one; the dilemma gym manufactures encounters on demand.
Why it ranks here. Simulation is the only way to give every learner volume — dozens of morally loaded reps a term rather than the handful life happens to supply — and the evidence says simulated situations train something text cannot. VR moral dilemmas elicit different responses than written vignettes, opening the gap between moral judgment and moral action to direct training (Francis et al. 2016; angle 06), with a current review mapping the method’s findings and limits (Europe PMC 2025; angle 06) and time-pressured VR showing decision behaviour as it actually runs (Sütfeld et al. 2017; angle 06). Structured dilemma discussion measurably shifts moral reasoning on the Defining Issues Test (ERIC 2015; angle 06), and the Konstanz Method provides the manualised, teacher-trainable delivery format (Lind 2021; Lind 2019; angle 06). Clinical simulation builds moral courage by rehearsal — the “Should I Say Something?” curriculum trains confronting professionalism lapses (Europe PMC 2023; angle 06). The military integrates ethical scenarios into high-intensity field exercises (Europe PMC 2014; angle 06), wargaming surfaces moral choice under uncertainty — and RAND’s 1950s political-military games show it has done so for seventy years (Emery/TNSR 2021; angle 06) — and serious role-play games teach applied ethics in technical domains (Springer 2020; angle 06).
Methods that fit. KMDD dilemma sessions as the weekly baseline — cheap, manualised, classroom-scale (Lind 2021; angle 06) · role-play and serious games for domain ethics (Springer 2020; angle 06) · VR reserved for scenarios where embodiment and time pressure are the point (Francis et al. 2016; Sütfeld et al. 2017; angle 06) · scenario-injected field exercises for professional formation (Europe PMC 2014; angle 06) · LLM role-play as the scaling layer — rehearsing conflict with generative agents and receiving feedback (Shaikh et al. 2023; Park et al. 2023; angle 14), under the governance rails Report 3 details, including auditing the moral biases of any model deployed as interlocutor (Abdulhai et al. 2023; angle 14).
Signals & KPIs. Dilemmas rehearsed per learner per term — the volume metric · moral-competence and DIT score growth against baseline (Lind 2019; angle 06; Thoma 2014; angle 03) · scenario-library coverage: every canon family represented at three difficulty levels · facilitator certification counts · transfer checks: observed behaviour in later unannounced situational assessments.
Institutional wiring & first moves. First moves: commission a national scenario library keyed to the 32 values, drawing on real anonymised cases from schools, hospitals, firms and public administration; certify a first cohort of KMDD-trained facilitators; pilot one LLM role-play deployment under human supervision with a published bias audit; resist the gravitational pull of buying VR hardware first — the method stack starts with a table and a dilemma.
Area 4 — The failure ladder: a designed progression of survivable failures
In short. For each value family, a deliberate, graded progression of survivable failures — from the ten-minute unsolvable problem to the failed pitch to the expedition that goes wrong — so that every learner fails frequently, safely, at a calibrated level, and always with a debrief. The ladder is the institutionalisation of Step 3 of the loop.
Why it ranks here. Because it is the loop’s most evidence-backed step and the one schooling is currently engineered to prevent. Generation-before-instruction produces deeper learning than instruction-first (Kapur 2015; angle 05), through activation, gap-awareness and encoding mechanisms (Kapur & Roll 2018; angle 05), robust even at micro-scale (Ziegler, Trninic & Kapur 2021; angle 05). Desirable difficulties improve retention and transfer precisely because they degrade immediate performance (Bjork & Bjork 2011; angle 05). Errorful learning with corrective feedback beats errorless learning (Metcalfe 2017; angle 05). Error-management training outperforms error-avoidant training in a 24-study meta-analysis, with the advantage concentrated exactly where character lives — adaptive transfer to novel situations (Keith & Frese 2008; angle 05). And the framing layer is proven at national scale: a brief growth-mindset intervention reframing failure raised achievement in a randomized experiment of over 12,000 students (Yeager et al. 2019; angle 05). A school that never lets learners fail is not protecting them; it is detraining Perseverance, Resilience and Humility — the entire Strength family.
Methods that fit. Productive-failure problem design as the classroom rung (Kapur 2015; angle 05) · error-management framing — “errors are informative, hunt them” — for skills training (Keith & Frese 2008; angle 05) · growth-mindset framing wrapped around every rung (Yeager et al. 2019; angle 05) · an explicit error climate in which errors are discussable, per the error-learning conditions literature (Augsburg 2016; angle 05) · expedition and enterprise rungs at the top, where failure carries real but bounded consequence.
Signals & KPIs. Failure-exposure rate: designed failures experienced per learner per term, by family · retry rate: share of failures followed by a re-attempt within the term · debrief coverage: designed failures with a completed coach debrief — target is all of them · error-climate survey scores · adaptive-transfer performance on novel unannounced problems.
Institutional wiring & first moves. First moves: rewrite assessment policy so first attempts are formally costless — no designed failure may enter a grade; publish the ladder to learners and parents so failure is legible as design rather than malpractice; train teachers in the discipline of withholding instruction briefly — the hardest retraining in this playbook; build the top rungs into Area 2’s expeditions and ventures so the ladder ends in reality.
Area 5 — The habit protocol: the daily reps
In short. The infrastructure of daily and weekly repetition that converts corrected behaviour into automatic disposition: WOOP plans, implementation intentions, habit-stacked routines and metacognitive review, personalised to each learner’s current formation plan. Small, boring, and the difference between a value performed and a value possessed.
Why it ranks here. The tools in this area carry some of the largest and most replicated effect sizes in the entire library. Implementation intentions: d = .65 across 94 studies (Gollwitzer & Sheeran 2006; angle 07). WOOP/MCII: roughly 60% more self-disciplined studying in a school experiment (Duckworth et al. 2011; angle 07), with replication improving report-card grades, attendance and conduct (Duckworth et al. 2013; angle 07). Habit mechanics — context cues, repetition, reward — are the settled account of how behaviour becomes automatic (Wood 2016; angle 07), with a practical protocol built on the finding that automaticity plateaus over roughly 66 days (Gardner, Lally & Wardle 2012; angle 07). Grit predicts retention across military, workplace, school and marriage samples beyond ability (Eskreis-Winkler et al. 2014; angle 07). The EEF’s metacognition guidance gives schools seven evidence-based recommendations for teaching the self-regulation layer (EEF 2018; angle 07). And deliberate practice supplies the structure for the skill-like values — with Macnamara’s caveat kept in view: practice hours alone explain a bounded share of variance, so the protocol engineers cues, feedback and context, not just repetition (Ericsson et al. 1993; Macnamara et al. 2014; angle 07).
Methods that fit. A daily WOOP rep at a fixed time (Duckworth et al. 2011; angle 07) · if-then plans written for each learner’s live situations — “if I am about to hand in work I know is copied, then…” (Gollwitzer & Sheeran 2006; angle 07) · habit stacking onto existing school routines per the Gardner protocol (Gardner, Lally & Wardle 2012; angle 07) · COM-B as the design checklist when a rep fails to stick — capability, opportunity or motivation (Michie et al. 2011; angle 07) · weekly metacognitive review per EEF guidance (EEF 2018; angle 07) · the gratitude and strengths micro-practices with RCT evidence behind them — three good things, signature-strengths use (Seligman et al. 2005; angle 10).
Signals & KPIs. Rep completion rates per learner per week · automaticity self-ratings trending up across the 66-day window (Gardner, Lally & Wardle 2012; angle 07) · WOOP plans written versus executed · context stability: share of reps anchored to a fixed cue · teacher-observed spontaneous use — the value appearing unprompted, which is the point.
Institutional wiring & first moves. First moves: install a ten-minute daily formation slot in the timetable — the same slot every day, because the slot is the context cue; train every teacher in WOOP and implementation intentions, which are learnable in hours; connect each learner’s reps to their current formation plan from Area 7 so the reps are personal, not generic; audit after one term with COM-B and fix the failed reps by redesign, not exhortation.
Area 6 — Mentorship infrastructure: the human transmission channel
In short. A built, standards-run system that puts a trained mentor beside every learner in formation — because values pass person-to-person, and the coach watching the rep is not a metaphor but a staffing requirement. Mentorship is infrastructure, with published standards, screening, matching and monitoring — not a goodwill scheme.
Why it ranks here. The evidence base is old, randomized and sobering in both directions. The landmark Big Brothers Big Sisters RCT showed community mentoring cut drug initiation, violence and truancy — the study that legitimised mentoring as policy (PPV 1995; angle 11). A national survey shows mentored youth do better on aspiration, school engagement and leadership — and that one in three young people grow up without any mentor at all (MENTOR 2014; angle 11). A multilevel meta-analysis shows even naturally occurring mentor relationships meaningfully improve academic, behavioural and socioemotional outcomes (Rhodes Lab 2018; angle 11). But the direction of the evidence is as important as its strength: targeted, problem-specific mentoring outperforms non-specific friendship models (Rhodes Lab 2020; angle 11); benefits vary by youth risk profile and matches need programme support to survive (MDRC 2013; angle 11); and the field’s own operating standards — recruitment, screening, training, matching, monitoring and support, and structured closure — exist precisely because unmanaged mentoring underdelivers (MENTOR 2015; angle 11). The exemplar evidence adds a design law: attainable exemplars motivate; distant saints demoralise (Frontiers/Han et al. 2017; angle 11; Han 2018; angle 03) — with witnessing moral excellence triggering elevation and prosocial contagion as the transmission mechanism (Sparks et al. 2019; angle 11).
Methods that fit. Cognitive apprenticeship as the mentor’s method — model, coach, scaffold, fade (Collins, Brown & Newman 1987; angle 11) · the Elements of Effective Practice as the non-negotiable operating standard (MENTOR 2015; angle 11) · goal-directed mentoring keyed to the learner’s formation plan rather than generic friendship (Rhodes Lab 2020; angle 11) · deliberate cultivation of natural mentors — teachers, coaches, employers — alongside formal matches (Rhodes Lab 2018; angle 11) · exemplar pedagogy using near-peer, attainable models (Frontiers/Han et al. 2017; angle 11) · workplace coaching methods for the adult tier, with honest acknowledgement of where that evidence is thin (PLOS/Grover & Furnham 2016; angle 11) · Carr’s caution kept on the wall: role-modelling is not imitation theatre, and hero-worship is not formation (Carr 2023; angle 02).
Signals & KPIs. Share of learners in formation with a mentor meeting the published standards · match duration and structured-closure compliance (MENTOR 2015; angle 11) · share of matches that are goal-specific rather than non-specific · mentor training hours completed before first match · mentee outcomes tracked against the formation plan, not against attendance.
Institutional wiring & first moves. First moves: adopt the Elements of Effective Practice as the national standard and fund only compliant programmes; recruit the first mentor corps from professions with formation cultures — veterans, nurses, master craftsmen, athletes — and train them in the debrief method of Area 4; build the near-peer layer by making final-year learners junior mentors, which trains Stewardship on the canon side while it staffs the system; close the one-in-three gap deliberately, highest-risk learners first (MDRC 2013; angle 11).
Area 7 — Purpose-matching and personalization: the right training for the right soul
In short. The layer that makes formation personal: a diagnostic of each learner’s strengths, values and emerging purpose, and an individual formation plan that routes them to the right situations, mentors and reps. A gymnasium does not put every body on the same programme; neither does this one.
Why it ranks here. Purpose is the strongest motivational engine in the library, and matching is the multiplier on everything upstream. Purpose — a stable intention to accomplish something meaningful to the self and of consequence beyond the self — develops in adolescence and can be deliberately cultivated (Damon, Menon & Bronk 2003; angle 12), with the field review documenting its links to health, happiness and achievement and identifying which interventions move it (Templeton/Bronk 2020; angle 12). A one-time self-transcendent-purpose intervention improved GPA and self-regulation on tedious tasks — purpose is a trainable performance lever, not a luxury (Yeager et al. 2014; angle 12), and recent evidence shows career calling driving learning engagement through hope and self-regulated learning (Europe PMC 2026; angle 12). On the matching side: the definitive person-environment fit meta-analysis shows values fit predicts satisfaction, commitment and performance (Kristof-Brown et al. 2005; angle 12); strengths-based education supplies five operating principles — measurement, individualization, networking, deliberate application, intentional development (Lopez & Louis 2009; angle 12) — on the back of the VIA strengths classification (VIA Institute 2019; angle 10). And personalization at system scale is feasible: RAND’s studies across 62 schools found personalized-learning schools outperforming comparators in maths and reading with effects growing over time (RAND 2015; angle 12), with the follow-on implementation study showing what tailoring actually changes in practice (RAND 2017; angle 12) and state-level policy levers already mapped (Aurora Institute 2016; iNACOL 2013; angle 12).
Methods that fit. A structured purpose interview at each major transition, in the Damon tradition (Damon, Menon & Bronk 2003; angle 12) · VIA-style strengths assessment feeding a strengths-based plan (VIA Institute 2019; angle 10; Lopez & Louis 2009; angle 12) · an individual formation plan naming the learner’s current focus values, reps, mentor and next proof event · fit-aware routing of placements — matching service, enterprise and expedition roles to strengths and calling (Kristof-Brown et al. 2005; angle 12) · AI-assisted personalization of scenarios and reps as the scaling layer, inside the equity guardrails the OECD has mapped (OECD 2024; angle 14) — detailed in Report 3.
Signals & KPIs. Share of learners with a current formation plan, reviewed each term · purpose-score growth on validated measures (Templeton/Bronk 2020; angle 12) · strengths-use frequency in placements · fit indices between placement roles and learner profiles · engagement trends in the learner’s hardest subjects — the Yeager test of purpose doing real work (Yeager et al. 2014; angle 12).
Institutional wiring & first moves. First moves: train form teachers and mentors to run purpose interviews at ages 14 and 18; issue every pilot-school learner a formation plan within the first term; wire the plan into Areas 2–6 so situations, ladders, reps and mentors all read from it; refuse the failure mode of personalization-as-software — the plan is a relationship with a document attached, not a dashboard.
Area 8 — Adult and professional formation: academies, charters and honest institutions
In short. Formation does not end at eighteen; the state’s own academies — military, police, medical, teaching, civil service — are where it must run deepest, because these professions hold coercive and fiduciary power. The models exist. So does the documented failure mode.
Why it ranks here. The most complete formation systems on earth are professional academies. West Point runs a 47-month immersive leader-development system whose explicit output is people who live and lead honourably (West Point 2025; angle 08), operationalised in the Gold Book’s honor education, developmental experiences and cadet character assessment (West Point 2019; angle 08). Medicine has professional identity formation — the evidenced process by which students come to think, feel and act as physicians (Sarraf-Yazdi et al. 2021; angle 08) — anchored by the Physician Charter’s canonical commitments (ABIM 2002; angle 08). Business education has the Ivey leader-character framework: eleven character dimensions with judgment at the core, taught as competence rather than compliance (Ivey/Crossan, Gandz & Seijts 2013; angle 08). Empirically, a study of 240+ junior British Army officers shows where values feature in professional life and where ethics training actually moves moral judgment (Jubilee Centre 2018; angle 08); a five-year multipronged evaluation shows an honor code shifting academic integrity (ERIC 2010; angle 08); engineering ethics education has been reviewed at every level, gaps included (Martin et al. 2021; angle 08). And then the warning that governs this whole area: Lying to Ourselves documented how an institution that piles impossible compliance demands on its officers teaches them to lie routinely — dishonesty as a learned survival skill inside a values-proclaiming organisation (Wong & Gerras 2015; angle 08), with the authors’ seven-year retrospective showing how slowly such cultures mend (Wong & Gerras 2022; angle 08). A formation system that demands the impossible is a dishonesty factory with a values statement.
Methods that fit. The academy pattern: multi-year immersion where living the standard is the curriculum (West Point 2025; angle 08) · professional identity formation staged across training, with socialisation treated as pedagogy (Sarraf-Yazdi et al. 2021; angle 08) · charters and codes as public commitments with teeth (ABIM 2002; ERIC 2010; angle 08) · leader-character frameworks embedded in professional education (Ivey/Crossan et al. 2013; angle 08) · ethics injected into field exercises rather than quarantined in classrooms (Europe PMC 2014; angle 06) · a standing requirements audit that finds and deletes impossible demands before they metastasise into learned dishonesty (Wong & Gerras 2015; angle 08).
Signals & KPIs. Professional-identity milestones assessed across the training arc (Sarraf-Yazdi et al. 2021; angle 08) · share of field exercises with embedded ethical scenarios · honesty-culture audits: count of requirements practitioners report as impossible to meet honestly, trending to zero (Wong & Gerras 2015; angle 08) · leader-character assessments validated against follower outcomes (Monzani, Seijts & Crossan 2021; angle 08) · honor-system data reviewed for learning, not for show trials (ERIC 2010; angle 08).
Institutional wiring & first moves. First moves: pick two academies as pilots and give each a Gold-Book-style operational character programme — named values, developmental experiences, assessment (West Point 2019; angle 08); write the profession’s charter where none exists, with the professional body, not the ministry, holding the pen (ABIM 2002; angle 08); run the first requirements audit within six months and publish what was deleted; make the audit annual and unkillable.
Measurement without Goodhart
Everything above generates data, and here the playbook must be most careful, because character measurement destroyed by its own stakes is worse than no measurement at all. The library’s measurement angle is blunt about the trap: self-reported character is confounded — by faking, and by reference bias, in which learners rate themselves against local peer standards, so the best schools can produce the worst-looking scores (RAND 2014; Heckman & Kautz 2014; angle 13). The KIPP evaluation is the canonical parable: rigorous methods found strong achievement effects and flat character-survey effects in an intensely character-focused network — very plausibly because the surveys measured shifting reference points rather than character (Mathematica 2015; angle 09). A state that builds league tables on such instruments will get gaming, not formation.
The answer is multi-method measurement, held deliberately low-stakes. Four instruments triangulate. First, situational judgment tests — realistic scenarios with scored responses — which resist faking better than Likert self-report and now have validated character-adjacent exemplars (Europe PMC/PLOS 2019; angle 13), alongside dilemma-based instruments fielded on 10,000+ UK students (Jubilee Centre 2015; angle 13) and the moral-competence and DIT traditions (Lind 2019; angle 06; Thoma 2014; angle 03). Second, observed behaviour: debrief records, retry rates, honor-system data, placement reviews by external users — behaviours over statements, the Heckman rule (Heckman & Kautz 2014; angle 13). Third, 360-degree perspective — mentor, peer and placement-supervisor ratings, in the multimethod spirit of assessments like ACT Mosaic (Europe PMC 2022; angle 13). Fourth, longitudinal outcomes: the point of the enterprise, tracked over years in the Heckman tradition. For the survey layer that remains, use the serious psychometric machinery the OECD has already built — sampling, scaling, invariance testing and anchoring vignettes that partially correct reference bias — in its international social-emotional skills survey (OECD SSES Technical Report 2025; angle 13; OECD 2021; angle 09). For values profiles, the revised Schwartz questionnaire measures 19 values with validated psychometrics across 49 cultural groups — as a mirror for the learner, never a grade (Schwartz & Cieciuch 2022; angle 13). The EASEL taxonomy keeps constructs comparable across programmes so the system does not drown in incommensurable vocabularies (Wallace/Harvard EASEL 2021; angle 13), and AIR’s readiness framework — stop, think, act — governs whether a school should be assessing at all yet (AIR 2015; angle 13).
Three rules make the system Goodhart-resistant, and they are constitutional, not advisory. Measure growth, not rank — every score is the learner against their own baseline; no league tables, ever. Never high-stakes — character data may not gate admission, employment or school funding; the moment it does, it stops being true (RAND 2014; angle 13). The learner owns the data — formation records belong to the person being formed, disclosed at their discretion; the institution keeps only anonymised aggregates for improving the gymnasium itself.
Governance guardrails
A state that trains values must answer the indoctrination charge before it is made, and the answer must be structural. The philosophical groundwork exists: Kristjánsson has systematically answered the ten standard objections to virtue-led character education — paternalism, relativism, situationism among them — and the core of the answer is that character education properly done trains judgment, the capacity to reason about the good, not obedience to a list (Kristjánsson 2013; angle 02). The pedagogy in this playbook is aligned with that defence by construction: dilemma discussion trains reasoning, not conclusions (Lind 2021; angle 06); phronesis — situational judgment — sits at the top of the canon (Jubilee Centre 2020; angle 02).
The comparative record supplies the cautions. Japan’s dōtoku shows a state moral curriculum can run for generations — and also how it accumulates political baggage and requires reform; the care-ethics and citizenship critiques of the post-reform subject are part of the record (NIER 2013; ERIC 2022; angle 15). Singapore’s CCE shows full systematisation is achievable — values, identity, relationships and choices architected from primary through secondary — and also that the architecture is state-defined top to bottom (Singapore MOE 2021; angle 15), a settlement a European democracy cannot simply import. Finland offers the alternative pole: a Bildung-based curriculum whose values core emphasises the learner’s own growth into humanity (Europe PMC 2021; angle 15), and the IB’s learner profile shows a values canon running successfully across borders under consent (ERIC 2025; angle 15). The synthesis for a state like the Czech Republic: a published, contestable canon; parliamentary and public consultation on its contents; explicit swap-rights for schools and communities within a common loop; parental transparency and genuine opt-outs; and judgment-training pedagogy as a statutory requirement. Add the humility the evidence demands — the largest randomized programme evaluation returned nulls (US IES 2010; angle 15), so every claim this system makes must be published with its measurement — and the honesty requirement of Area 8: an institution that preaches the canon while demanding the impossible will teach lying at scale (Wong & Gerras 2015; angle 08). Transparency of the canon, of the methods, and of the results is not public relations. It is the licence to operate.
The first twelve months — a Czech sequence
A mid-sized EU state can stand this up inside a year. The Czech Republic — ten million people, a respected school system, universal military and police academies, a live public debate about institutional trust — is the worked example.
Months 0–2: name the canon. Convene a standing formation commission — educators, philosophers, professional bodies, parents, and learners — and publish the draft 32-value canon for public consultation, ENSI’s version as the starting text, explicitly swap-ready. Adopt version one by decree of process, not consensus of everyone: the commission decides, publishes its reasoning, and schedules the first revision for year three. The naming is the founding act; do it in public.
Months 1–3: pick the pilots. Twenty schools — primary and secondary, Prague and Brno alongside small-town and rural, selected for willing heads rather than existing excellence, because Area 1 begins with leadership. Plus two professional academies — the military university and a medical or teaching faculty — as the Area 8 pilots, each committed to a Gold-Book-style programme and the first requirements audit (West Point 2019; Wong & Gerras 2015; angle 08).
Months 2–5: train the people. Certify the first facilitator cohort in dilemma method (Lind 2021; angle 06); train pilot-school staff in WOOP, implementation intentions and metacognitive routines (Duckworth et al. 2011; Gollwitzer & Sheeran 2006; EEF 2018; angle 07); stand up the mentor corps under the Elements of Effective Practice as the binding national standard (MENTOR 2015; angle 11).
Months 3–6: stand up the dilemma-gym MVP. A national scenario library covering all eight families at three difficulty levels, built from real anonymised Czech cases; weekly KMDD sessions in every pilot school; one supervised LLM role-play pilot with a published bias audit (Shaikh et al. 2023; Abdulhai et al. 2023; angle 14).
Months 4–7: baseline. Run the full multi-method baseline in every pilot: SSES-aligned survey instruments with anchoring vignettes (OECD SSES Technical Report 2025; angle 13), situational judgment and dilemma instruments (Europe PMC/PLOS 2019; Jubilee Centre 2015; angle 13), observed-behaviour rubrics, school-climate measures — and the purpose interviews that seed each learner’s formation plan (Damon, Menon & Bronk 2003; angle 12). No school proceeds unbaselined; growth is the only score this system keeps.
Months 6–12: run the loop. First failure ladders live in classrooms; first service placements and expeditions under Area 2 agreements with municipalities and NGOs; daily formation slots running the habit protocol; mentors matched to the highest-need learners first (MDRC 2013; angle 11). Month 12: publish everything — canon, methods, instruments, baseline data, costs, failures — openly, and invite EU peers to fork it. The publication is itself a rep: the state performing Honesty, Accountability and Stewardship in public, once, correctly.
Close: build the gymnasium
The evidence in this library converges on one uncomfortable sentence: every institution is already a value gymnasium — most are just running a bad programme by accident. The school that never lets a child fail is training fragility. The ministry that demands the impossible is training dishonesty (Wong & Gerras 2015; angle 08). The timetable with no service, no expedition and no dilemma in it is training the quiet conviction that values are things adults say, not things people do. Formation is not optional. Only deliberate formation is.
The deliberate version is now buildable, and nothing in it is speculative. The canon is a decision. The loop is assembled from mechanisms with meta-analyses behind them — productive failure, error management, implementation intentions, service-learning, mentoring, dilemma discussion. The programme areas have named exemplars running today, from a 47-month academy (West Point 2025; angle 08) to a 213-programme evidence base (Durlak et al. 2011; angle 09) to national curricula already operating at state scale (Singapore MOE 2021; angle 15). The measurement exists, with its traps mapped (Heckman & Kautz 2014; angle 13). The AI layer that makes situational practice personal and abundant is the subject of Report 3.
What remains is the founding act, and it costs almost nothing: name the canon, in public, and put the first twenty schools on the floor. A state that starts now buys the option on everything downstream — the trust dividend, the honesty dividend, the compounding advantage of a generation trained by reps rather than exhortation, visible in the data within a decade because childhood formation predicts adult flourishing better than grades do (LSE CEP 2013; angle 10). A state that waits will keep delivering lectures on strength to citizens it never once put under load. The gymnasium is the oldest idea in education, and it has been waiting twenty-four centuries for an operator. Open the doors.




