ENSI — European Nexus for Strategic Intelligence. The companion to “Civic Apps: Value for Society”, which sets out the principles behind these plays.
A town council argues for three hours about whether to build a new road. The session is streamed and recorded, as the law requires. Almost nobody watches it live, and nobody watches the recording afterwards. A week later the local paper runs four paragraphs, a resident posts an angry summary of what she thinks was said, and the argument that actually took place — who proposed what, which facts were disputed, where the councillors quietly agreed — is lost. Democracy happened in public and nobody could see it.
The software to change that already exists. An app can join the session, build a map of the argument as it unfolds — every claim tied to the words that were spoken — and publish it next to the recording, so that anyone can follow three hours in ten minutes and check exactly who said what. It was built by a very small team. It is not yet used in any council chamber.
That is the situation of almost every civic app in this body of work. They are built, or nearly built, and they are stuck — not because the code is unfinished but because the last mile runs through things code cannot produce. The companion piece, Civic Apps: Value for Society, names those things as five barriers: data that does not exist or is not released, no institution that owns the app, no person who signs what it says, no path to the people who need it, and no money that does not come with strings.
This piece is about getting past them. It sets out eight plays — eight concrete moves, each carrying specific apps across the last mile to a place where they change something for people outside the team, and where that change can be shown. Each play names the apps it builds on, describes the idea and how it would work in practice, explains why it sits where it does in the order, says how you would know it worked, and says when to stop.
The order matters. Each play was weighed on four questions: how much public value it could create, how ready its apps already are, how much pull there is from an institution that wants it now, and how much risk it carries if it goes wrong. The plays are ranked by how soon each can produce value that someone outside the team can see — not by which has the largest theoretical upside. The first plays are the ones where a partner is waiting and the harm of a mistake is small; the later ones are the most important and the most dangerous.
None of the plays is a new product. Every one of them is a way of finishing something that already exists.
The eight plays in brief
Turn the state map into a crisis picture — give the people who manage emergencies one view of what depends on a failing system, who is affected, who is responsible, and what they decided.
Offer evidence as a public service — answer the questions a parliamentary committee is actually facing with verified evidence packs, published for everyone.
Map public deliberation — publish cited maps of council sessions, hearings and debates so any citizen can follow and check what was said.
Form the builders — create a pathway that turns motivated people into builders and stewards of governed agent systems.
Carry health literacy into under-served languages — research, translate and clinically review health knowledge for the people the market ignores.
Calibrate the priority register — check the AI that ranks where talent and money do most good against human experts before anyone relies on it.
Build accountability with people in charge — verify public claims at machine speed but publish at human pace, with a named editor, a right of reply and prominent corrections.
Take the guardrails to the rule-setters — put the missing AI standards and the companion-AI charter in front of the institutions that can adopt them.
1. Turn the state map into a crisis picture
This play builds on Mapa státu, the Resilience map, Holoprism‘s scouts and sentinel, and MeetViz sitting in the crisis-staff meeting. Its idea is simple to state: when a critical system fails, the people managing the crisis should see, in one place, what depends on that system, which towns and services are affected, who is responsible for each part, what signals are coming in — and a cited record of the decisions they are taking as they take them.
In practice it would work like this. A system goes down in the middle of the night. The duty officer opens the map of the state and switches off the failed system; the dependency cascade spreads across the screen to first- and second-order effects, with the municipalities, services and population affected and the authorities responsible for each. Scouts pull in the signals that matter — outage reports, operator notices, public statements — and a sentinel flags when a second dependency starts to wobble. When the crisis staff meet, a bot joins the call and builds a live map of what is being decided, so that the record of who agreed to what exists the moment the meeting ends. The morning briefing is assembled from that record, not from memory.
It comes first because it is the closest to being real. The state map already exists and has been reshaped, version after version, by the working group that uses it, and a crisis-management community already wants exactly this. Most of the data is public. The one missing layer — who actually operates which system for which authority — is precisely what that community is positioned to request from the state, which means the same partner clears both the data barrier and the mandate barrier.
You would know it worked not from traffic but from outcomes. The first is failure points that were identified and then fixed by the bodies that own them — a single point of failure nobody knew about, discovered on the map and resolved. The second is time: in a tabletop exercise, how many minutes it takes a crisis staff to agree on the picture of the impact with the tool, and how many without it.
The stopping rule is just as clear. If no public body will supply or confirm the missing operator data within a year, the crisis features should not be built on guesses. The map stays a research instrument, valuable for analysis but not trusted for action. The first move is a single tabletop exercise with the working group, run on a realistic scenario, with the timings recorded.
2. Offer evidence as a public service
This play builds on the ENSI research engine and Hyperthesis‘s claim audit. Its idea is to turn what these engines already do — produce sourced, checked research on any question in days — into a standing public service for the people who write laws. A parliamentary committee or research office sends a question it is actually facing. It gets back a verified evidence pack: every figure checked against its source, every gap declared, every contested point shown as contested, and the method written down. The pack is published for everyone by default.
In practice, a committee preparing a bill on, say, energy security sends three questions a fortnight before its hearing. The engine plans the angles, fetches the primary sources, drafts the pack section by section, and a separate layer of critics attacks it for missing counter-evidence, unsupported figures and stale sources. A named editor reviews it and signs. The committee receives it in time to use it, and the public can read the same pack the legislators read — including a list of what could not be found.
It sits second because the engines already work and their output is evidence rather than verdicts, which keeps the risk of harm low. A single partner who signs clears both the mandate barrier and the distribution barrier: once a committee cites a pack in its papers, the pack has found its audience and its purpose. And the discipline the engines already follow — never cite what was not fetched, declare what could not be obtained — is exactly what a legislative audience needs before it can trust material assembled by agents.
It worked if packs are cited in committee papers, amendments or debates, and if staff say they would otherwise have had to commission the work or go without. The value can be estimated honestly as the research that would otherwise have been bought, reduced by the share that was actually used.
If none of the first half-dozen packs is used, the right response is to change the audience, not the engine — some offices are ready for this kind of service and some are not. The first move is to offer a committee three pilot packs on questions already on its agenda.
3. Map public deliberation
This play builds on MeetViz and Noeverse. Its idea is the one from the opening of this piece: publish cited maps of the places where public decisions are argued in public — council sessions, parliamentary hearings, televised debates — next to the recording, so that a citizen can follow two or three hours of argument in ten minutes and check, claim by claim, what was actually said and by whom.
In practice, the app joins a council session that is already being streamed. As the session unfolds it builds a map of the argument: the proposals, the claims made for and against them, the facts that were disputed, the points on which councillors who usually disagree turned out to agree. Every node in the map links to the moment in the recording where the words were spoken. When the session ends, the map is published next to the recording, and anyone can ask it a question and get an answer that points back to the source.
It sits third because it needs almost no new building and carries little risk. The output is a map, not a verdict, and every part of it points to the words spoken, so a mistake is visible and correctable. Its reach rides on meetings that already happen and recordings that already exist, which quietly solves the distribution barrier. And it puts one of the deepest things these apps can do — turning a stream of talk into a structure people can see — to the most public purpose there is.
The test is comprehension. Give one group of residents the map and another the official minutes, and ask both what was decided and why. If the map does not help people understand the session better, it is decoration, however beautiful. A second measure is simpler: whether residents, journalists and councillors themselves start linking to the maps when they argue about what happened.
The play should stop if citation accuracy ever falls below the apps’ own quality bar, because a map that misattributes words in public does more harm than no map at all. The first move is one municipality that already streams its sessions, one season of published maps, and one comprehension test.
4. Form the builders
This play builds on Aether University, the Agent-Driven Meetup and Machine Game. Its idea is a pathway that turns motivated people into builders and stewards of governed agent systems: courses with agent-curated reading, the live critique of real ideas in a room of peers, and a game in which you learn to govern a company run by AI employees before you are trusted with a real one.
In practice a cohort of fellows — some from companies, some from public institutions, some starting out — works through a sequence. They read primary texts chosen and assembled by agents, argue about them in seminars, and learn to state intent precisely enough for an agent fleet to act on it. They bring their own ideas to the meetup format and watch them critiqued live, by peers and by agents. They play through the governance of an AI-run organisation and discover, safely, what happens when approval levels are set wrong. And each of them ends by carrying a real civic app one step further — which is how the pathway feeds every other play on this list.
It belongs here because AI’s leverage flows to the people who can direct it, and the most direct way to spread the gains is to widen who can. Every other play depends on people who can specify intent, judge an agent’s output and govern a fleet; without them, civic apps remain a handful of builders’ side projects. Formation is also the play in which the builders’ own experience transfers most directly: what was learned building these apps is exactly what the next generation needs to learn.
It worked if graduates build and steward things that a similar group who did not take part does not — measured honestly, with a comparison group rather than testimonials. A second measure is who takes part: a pathway that only forms people who were already advantaged has not widened anything.
If after two cohorts the graduates do no better than the comparison group, the programme should be redesigned before it grows. The first move is a single cohort with a partner university or employer, with the curriculum and the outcome data published.
5. Carry health literacy into under-served languages
This play builds on audhd.cz and Believer‘s grounded-evidence pattern. Its idea is deep research plus terminology-exact translation, grounded only in real sources, for conditions and languages the market neglects — reviewed by clinicians before anything is published, and written for the people who live with the condition rather than for the professionals who study it.
In practice, a research agent works through the international literature on a condition, phenomenon by phenomenon and perspective by perspective. A translation step renders it into the target language using a controlled glossary, so that the medical terms are the ones clinicians in that country actually use. A grounding rule allows the text to cite only real studies and real people, never invented ones. A clinician reviews every piece before publication and signs off. The result is a body of accurate, readable material in a language that no medical publisher would have found worth the cost.
It is a strong play because it is cheap, measurable and a clear case where AI turns an impossible cost into a small one. A small language community has always been too small to justify a medical publisher’s attention; it is not too small for a pipeline that can research, translate and check at almost no marginal cost. The value is immediate and personal — a parent who finally understands their child, an adult who finally recognises themselves.
It worked if the material reaches people in their language, if clinicians confirm its accuracy, and if readers report, in a simple opt-in survey, that they understand their situation better. The stopping rule is the strictest of any play: if anything clinically unsafe is ever published, the pipeline pauses until review is fixed. The first move is a clinical partner to review the existing material and a patient organisation to distribute it.
6. Calibrate the priority register
This play builds on Hard Frontier, Climate Lab and Future of Politics. Its idea is one transparent register of where talent and money would do the most good — evidence shared, weights your own — in which anyone can see why a task, a technology or an issue ranks where it does, and change the ranking by bringing their own values. The decisive step is to check the AI raters against a panel of human experts before anyone relies on the rankings.
In practice the register brings together three things that already exist: scores of how much better the world becomes if a capable person spends a year on a task, scores of climate technologies on many dimensions that users can re-weight, and a list of the decisive issues that no institution owns. Each score comes with its reasoning and the evidence behind it. A funder can apply their own weights and see how their priorities change the ranking; a student can see which kinds of work matter most by their own lights. Before any of it is presented as a ranking, a panel of human experts scores a sample of the same items independently, and the agreement is published.
It sits in the middle of the list because its value is high if it is trusted and close to zero if it is not, and trust depends entirely on calibration. Two independent AI raters agreeing with each other is encouraging; AI raters agreeing with people who have spent their careers in the field is evidence. Calibration is a finite, fundable piece of work, and it is what turns a persuasive instrument into a credible one.
It worked if the agreement between AI raters and experts is good enough to publish without embarrassment, and if money or careers are actually redirected because of the register. If the agreement stays low even after calibration, the honest move is to publish that finding and keep the register as a tool for discussion rather than a ranking. The first move is a small expert panel for one domain — climate or AI safety — and a published calibration study.
7. Build accountability with people in charge
This play builds on DemocracyWatch and Bohemie. Its idea is to verify public claims at machine speed but to publish at human pace: a named editor signs every verdict, the person checked gets a right of reply before publication, corrections are as prominent as the original findings — and what goes right in the country is measured as rigorously as what goes wrong, so that the whole effort cannot be dismissed as one side’s weapon.
In practice, agents harvest public statements from debates, interviews and posts, extract the checkable claims, and verify each one in a bounded number of research rounds, stopping as soon as independent sources settle it. The result is a proposed verdict with its evidence attached. A human editor reviews it, the person whose claim it is gets the chance to respond, and only then is anything published — under the editor’s name. Alongside the verdicts, the same machinery measures democratic health on many dimensions and tracks the good news that ordinary coverage misses, and every measurement ends in an action a citizen, a journalist or an official could take.
It sits seventh because it carries both the largest democratic upside and the largest risk. It needs real data pipelines rather than placeholder numbers, a publisher willing to sign, and trust from people across the political spectrum — and all three take time to earn. The fastest route is not to compete with existing fact-checkers but to partner with one, offering machine-speed verification as support for its editors rather than as a replacement for them.
It worked if the people checked start correcting themselves, if trust in the verdicts holds across political camps, and if the error rate is published and low. The stopping rule is unforgiving: if a false verdict about a named person goes uncorrected for a day, automated proposals stop until the process is fixed. The first move is a partnership with one established fact-checking organisation, starting as a verification assistant for its existing editors.
8. Take the guardrails to the rule-setters
This play builds on AGI Standards, Agentic Safety and the Sentience Accord. Its idea is to put three pieces of work that already exist — a map of the measurable AI standards that are missing and who should own each, an engineering reference for keeping autonomous agents safe, and a charter for how companion AI must treat the people who rely on it — in front of the standards bodies, regulators and labs that could actually adopt them.
In practice this means following the consultations, working groups and drafting processes where the rules for AI are being written, and submitting the relevant parts of the work at the right moment: the missing standard to the committee drafting that area, the safety reference to the guidance on agents, the charter’s provisions to the people shaping rules for companion products. Agents can do the watching — tracking consultations and drafts, mapping each against the gap list, preparing submissions — while people do the arguing.
It comes last not because it matters least but because its value depends entirely on others. The work is already written and it costs almost nothing to submit; what it needs is persistence, patience and a willingness to let institutions take ownership of ideas that began elsewhere.
It worked if provisions are cited or adopted in standards, guidance or law. There is no stopping rule, because the cost is so small — only the obligation to report honestly what was, and was not, taken up. The first move is a submission to the next relevant AI Act implementation consultation and to one national standards committee.
The first twelve months
The plays are not meant to run all at once. A sensible year sequences them so that the cheapest evidence comes first and the riskiest commitments come last.
In the first quarter, measure before building anything new. Start a value ledger for every civic app, even if at first it records only effort, outputs and use. Write a one-page statement of the outcome each of the first three plays is meant to cause. Remove, or label unmistakably, every simulated or illustrative number in anything public. And find the institutional owner for the most-used civic app.
In the second quarter, run the first three plays with real partners: a tabletop crisis exercise using the state map as a crisis picture, a small set of evidence packs for one committee or research office, and a season of deliberation maps for one municipality, each with its simple test — minutes saved, packs cited, comprehension compared.
In the third quarter, turn to people and evidence: run one formation cohort with a comparison group, put the health-literacy material through clinical review, and recruit the expert panel that will calibrate the priority register.
In the fourth quarter, publish and decide. Publish the value ledger and an honest account of the first three plays, including what faded and what would have happened anyway. Begin the accountability play only if a fact-checking partner has signed and the human-signature process is in place. Submit the guardrail work to at least one consultation and one standards committee. Then choose the next year’s plays by what the ledger shows, not by what is most exciting to build.
Choosing the first play
Every play on this list is a way of finishing something rather than starting something, and that is the point. The civic apps do not need more features; they need partners, permissions, signatures, audiences and evidence. Each play is a route to those things for a particular app.
If only one play can be run, run the first: it has a partner waiting, a missing piece that the partner can supply, and an outcome — a failure point found and fixed — that anyone can understand. If three can be run, run the first three, because together they touch the state, the parliament and the public, and they carry the least risk of harm. The later plays are more important in the long run and should be prepared now, but started only when the trust, the data and the partners they depend on are in place.
The measure of success for the whole programme is simple. A year from now, can someone outside the team — a crisis officer, a committee researcher, a resident of a town — point to something that changed because a civic app existed? If yes, the apps have started to create value for society. If not, they are still only software.
This piece draws on the Metamatics Ventures and ENSI NGO apps and on ENSI’s research library on public value, digital public infrastructure, civic technology and AI governance.



