The architecture, the agents, the loops and the principles.
ENSI — European Nexus for Strategic Intelligence. Companion to “Sixteen Principles”. Compiled 2026-09-19, revised 2026-09-21.
The argument, before the list
When building becomes cheap, a small team can start many things: a sales-agent workforce, a companion-AI charter, a meeting mapper, a health-literacy site, a map of the state. Read one at a time, such work looks like a pile. Read as a whole, something else appears. Together these kinds of software form the organs of a society’s capacity to govern itself — a way to see what the state is, to know what is true, to make sense of what is being said, to decide what matters next, to hold power to account, to set rules for the agents doing the work, an economy that pays for it, and the people and communities who carry it.
Each organ can be built separately, and usually is. The architecture is the same each time: an orchestrator, a named fleet of specialist agents, tools reached through fallbacks, a deterministic validator around every model call, a human gate, an audit log, cheap models for volume and strong ones for judgement, and an agent-facing surface. That architecture is proven first in market products, where customers can say no. It is used least in public ones.
This report argues that the “greater thing” is not another product. It is the composition: a Civic Agent Stack in which the twelve contribution types become modules on a shared substrate, reading and writing three shared graphs — of the state, of the evidence, of public discourse — and running four loops that every healthy democracy runs today slowly and mostly in people’s heads: see → decide → act and account → learn and become. A mid-sized European democracy, such as the Czech Republic, is the natural first home: large enough to matter, small enough for a few institutions to adopt the stack.
Three principles drive the design. Reuse beats invention: the stack’s components are the same few patterns every serious executor product ends up building; build them once and share them. The executor must enter public life under public rules: civic modules should be the most governed agents in the stack — human gates at every decision point, provenance on every claim, published audit trails. The market pays for the commons: commercial products field-test and finance the patterns; a public-interest organisation opens and deploys them. That division of labour is how a very small team can build at the scale of an institution without becoming one.
The main points
The parts already exist as patterns. Every module a civic operating system needs corresponds to a contribution type that AI-native builders already produce; the gap is integration and public deployment, not invention.
Six shared components form the substrate: a governed tool dispatcher, a validator library, an evidence ledger with provenance, a tiered model router, a skills registry and an agent-facing protocol surface.
Three graphs form the data spine: the State Graph (what the state is), the Evidence Graph (what is known) and the Discourse Graph (what is being said) — plus a consent layer for personal context.
A standing fleet of twelve civic agents — Scout, Librarian, Cartographer, Analyst, Verifier, Adversary, Forecaster, Steward, Sentinel, Herald, Tutor, Warden — covers four loops; only the few that make consequential decisions do so, and never without a human.
Five composite capabilities are greater than the sum: a crisis common operational picture, a public evidence service for legislators, a national priority register, a civic accountability desk and a builders’ formation pipeline. No single module can deliver any of them.
The operating model is a studio with two sides: commercial products field-test and fund the patterns; a public-interest organisation opens and deploys them. Design for the broad-gains branch of the “regime fork”.
Governance is the product: a published roster of agents, gated decisions, visible activity, safe defaults and measured honesty — the civic agents are the most governed software in the stack.
Twelve months, four quarters: extract the substrate; stand up the graphs; ship two composite capabilities with partner institutions; publish the audit and measure the value.
How this report is organised
The stack is specified from the bottom up: the substrate (§1), the data spine (§2), the fleet of civic agents by archetype (§3), the four loops they run (§4) and the composite capabilities that emerge (§5). Then the operating model (§6), governance (§7), a twelve-month build (§8), the risks (§9) and a close.
1 · The substrate — six components every executor product rebuilds
The most expensive mistake a builder of many agentic products can make is to rebuild the same infrastructure in each one. Validators, dispatchers, ledgers and model routers tend to be written anew, product after product. The first move towards something greater is to extract six components into a shared substrate every module — commercial or civic — sits on.
1.1 The governed dispatcher
What it is. One choke point through which every tool call and external action passes: policy check → approval gate → cost estimate → loop detection → credentials → execute → trust envelope → verify → log → emit — with approval levels (none, notify, approve, maximum) set per activity, and per-agent tool allow-lists, deny-lists and “never” rules.
Why the stack needs it. It is where public rules become enforceable. A civic agent that publishes, queries a state system or contacts a citizen does so through the dispatcher, under a policy the public can read.
In the apps: Kybernetist’s single gate for every tool call and its per-activity approval levels; Primespect’s per-agent tool lists and “never” rules.
1.2 The validator library
What it is. Shared deterministic checks that wrap model outputs: schema validation and repair; citation and quote validation; query safety (read-only, single statement, bounded, pre-checked); substring provenance (every claimed source span must literally exist in the source); redundancy detection; language detection before model calls; verification that blocks publication.
Why the stack needs it. “Never trust the model’s own verdict” is the most transferable principle of the executor era. In public life it is not good practice but a duty.
In the apps: the state map’s query-safety check, Holoprism’s quote check, MeetViz’s citation check, Primespect’s qualification gate, the Sentience Accord’s verify-before-publish.
1.3 The evidence ledger
What it is. An append-only record of every claim the system holds — source document, quoted span, confidence, the agent that produced it, the human who approved it — with a provenance view on every screen, and a declared list of what could not be obtained.
Why the stack needs it. Public trust in agent-produced knowledge rests on anyone being able to ask “where did this come from?” and get an answer that does not depend on trusting the agent.
In the apps: Holoprism’s evidence ledger and provenance drawer; Scaffold’s constitution for claims; the research engine’s declared “missing” list.
1.4 The tiered model router
What it is. A service that gives each job the cheapest adequate model, escalates by consequence, records per-call cost, and routes by region for latency and data residency — with a zero-model deterministic tier wherever code can do the job.
Why the stack needs it. A civic stack runs millions of small judgements. Its economics and its European data-residency obligations are both decided here.
In the apps: Primespect’s tiering by the cost of a mistake; Holoprism’s zero-model harvest tier; the state map’s EU-region model for data residency.
1.5 The skills registry
What it is. Institutional procedures written as versioned playbooks that agents execute and people can read, diff, challenge and improve.
Why the stack needs it. It is how a method becomes public: the procedure by which a fact-check, a foresight scan or a resilience assessment is done is itself published and open to challenge.
In the apps: Complexity’s executable business skills; the ENSI research method; Aether’s reading-selection skill.
1.6 The agent-facing surface
What it is. Every module exposes its capabilities as tools other agents can call — with identity, scopes, audit and revocation.
Why the stack needs it. Citizens, journalists and officials will increasingly reach public information through their own agents. A civic stack that exposes cited, governed tools becomes the grounding layer those agents use — instead of whatever their training data happens to remember.
In the apps: Primespect’s whole API as agent tools; Holotext’s scoped, revocable context tools; the WHY Engine.
2 · The data spine — three graphs and a consent layer
Modules become a system when they share data. Three graphs carry the stack.
2.1 The State Graph
What it holds. What the state is: every agenda or area of public competence, every authority, every registered information system, every service, every data object and permission — and the edges among them. Many European states already publish much of this as open data in their registers of public administration.
What it enables. Browsing and querying the administration in plain language; simulating outage cascades to chosen depth; seeing which municipalities, services and people depend on a single system.
In the apps: Mapa státu and the Resilience cascade map.
What is usually missing. The operator → system → authority layer (who actually runs what, for whom), and live feeds rather than dated snapshots. Both require the state as a partner.
2.2 The Evidence Graph
What it holds. What is known: documents, the claims extracted from them, and the causal edges between factors — each with provenance, confidence and a verdict.
What it enables. Answering a question with cited evidence; seeing which claims rest on one source; tracing second-order dependencies (”what influences, depends on, regulates what”) with an adversarial check on every causal edge.
In the apps: the research libraries, Scaffold’s cited claims, Dependence Mapper’s causal edges.
What is usually missing. Extraction. Libraries of downloaded documents are records waiting for executors; the graph appears only when claims and edges are extracted, verified and linked.
2.3 The Discourse Graph
What it holds. What is being said in public: claims made in debates, hearings, interviews, meetings and statements; their argumentative structure; their verification status.
What it enables. Cited maps of public argument; verification of claims in bounded rounds; bridging views of where people who usually disagree actually agree.
In the apps: Noeverse’s claim graphs, MeetViz’s meeting maps, Lhari’s verification design.
What is usually missing. A shared claim identity — so the same claim, made in a parliamentary debate and repeated in a podcast, is one node with one verification history — and links to the State and Evidence Graphs.
2.4 The consent layer
Personal context — a citizen’s or an official’s own documents, mail, notes — is not a public graph and must never become one. The right primitive for joining it to public work on the owner’s terms is consent-based context lending: scoped, redacted, revocable and audited. An official preparing for a crisis meeting could lend the stack a slice of their briefing notes for one question and revoke it afterwards — the model Holotext already implements.
2.5 Where the value is: the edges between graphs
A public claim about hospital waiting times (Discourse) is verified against documents (Evidence) and linked to the agendas and authorities responsible (State). A foresight signal about a cloud provider’s outage (Evidence) is linked to every public system that depends on it (State) and to officials’ statements about digital sovereignty (Discourse). A parliamentary question is answered with cited evidence, the responsible authority and the cascade risk in one view. No single module can produce that answer; the spine can.
3 · The fleet — twelve standing civic agents
Agentic products tend to accumulate dozens of named agents, many doing the same job under different names. A civic stack needs a small, stable, publicly documented roster — agents whose roles, tools, limits and “never” rules any citizen can read. Twelve are enough, grouped by archetype, because the archetype decides how much autonomy each may have.
Read-only assistants — answer, never act
1. The Librarian answers questions over the three graphs with citations, through validated queries only. Never: writes, publishes, or answers without a source.
2. The Tutor explains — a regulation, a budget line, a scientific claim — at the reader’s level, tagging every statement as established, likely, contested or speculative, with Socratic scaffolding rather than answers on demand (unguarded AI tutoring makes learners worse). Never: presents a contested claim as settled.
Workflow executors — run fixed procedures; humans approve outputs
3. The Cartographer turns long material into structured maps: a debate into a claim graph, a meeting into a cited map, a paper into ideas and mechanisms. Never: keeps a node without a verbatim citation.
4. The Analyst runs published procedures — a resilience assessment, a regulatory-impact check, a priority scoring — as versioned skills. Never: runs an unpublished procedure.
5. The Herald assembles briefings for named audiences — a crisis staff, a committee, a newsroom — from approved material only. Never: sends anything external without a human signature.
Autonomous workers — act within hard bounds
6. The Scout watches sources continuously — gazettes, registers, feeds, publications, news — and harvests what changed, with a zero-model tier first. Never: bypasses a source’s access rules or keeps a file it could not validate.
7. The Verifier checks claims against evidence in bounded rounds (capped rounds and tokens, stopping when independent sources refute, writing only from fetched notes) and proposes a verdict. Never: publishes the verdict.
8. The Sentinel raises early warnings when a monitored indicator crosses a threshold or a dependency breaks. Never: acts on the warning itself.
Decision-makers — gated, logged, reversible
9. The Steward is the only agent that may move something from proposed to published — a verdict, a score, a correction, a change to the graphs — and only with a named human’s approval recorded in the ledger. Never: acts without a human signature.
10. The Warden calibrates and, if needed, suspends other agents, comparing their outputs against human judgements and outcomes and enforcing a one-way quality ratchet. Never: is itself unaudited.
Multi-agent systems — research and foresight
11. The Forecaster is an ensemble: several models and methods produce forecasts and scenarios, which are aggregated, red-teamed and scored against what happens. Never: reports a point forecast without its range and its track record.
12. The Adversary attacks everything the others produce — reversed causality in an edge, a missing counter-case in a report, a leading question in a briefing. Never: is switched off to meet a deadline.
The design principle: autonomy where the task is reading and structuring; human signatures where the output changes what the public is told.
4 · The four loops
A self-governing society runs four loops whether it knows it or not. The stack runs each continuously, with humans at the points of judgement.
Loop 1 — See
Question: what is true, what is changing, and what is being said about it? Agents: Scout, Librarian, Cartographer, Adversary. Contribution types: Evidence Engine, Sense-Making Instrument, State X-Ray. How it runs: scouts harvest changes; the cartographer structures them; the Evidence and Discourse Graphs update with provenance; the adversary attacks new links; the librarian answers anyone’s question with citations. Human role: editors curate sources and approve schema changes.
Loop 2 — Decide
Question: what matters most, what is likely to happen, and what are the options? Agents: Forecaster, Analyst, Adversary, Herald. Contribution types: Priority & Foresight Engine, Evidence Engine. How it runs: the forecaster ensemble produces scenarios and probabilities; the analyst runs published scoring procedures; the adversary red-teams; the herald briefs a named decision-maker. Human role: decision-makers decide. The stack makes options, evidence and uncertainty legible — it never chooses.
Loop 3 — Act and account
Question: who is responsible, did they do what they said, and what broke? Agents: Verifier, Sentinel, Steward, Librarian. Contribution types: State X-Ray, Accountability Engine, Guardrail. How it runs: claims and commitments are linked to the responsible authorities in the State Graph; the verifier checks claims in bounded rounds; the sentinel watches for broken dependencies; the steward publishes verified findings only with a human signature; corrections are logged as publicly as the original. Human role: named editors own every published verdict; affected parties get a right of reply before publication.
Loop 4 — Learn and become
Question: are people becoming more capable of seeing, deciding and acting well? Agents: Tutor, Herald, Cartographer. Contribution types: Formation Engine, Wellbeing Companion, Movement & Commons Layer. How it runs: what the other loops produce becomes learning material — curated reading spines, simulations of governing an agent-run organisation, live critique formats, health literacy in under-served languages — with a tutor that explains at calibrated certainty. Human role: teachers and community organisers. The loop’s measure is capability formed, not content delivered.
5 · Five capabilities greater than the sum
The case for composition rests on capabilities no single module can deliver, however good.
5.1 A crisis common operational picture
The need. When a critical system fails, crisis staff need to know within minutes what depends on it, who is affected, who is responsible and what is being said. The apps: Mapa státu, Resilience, Holoprism, MeetViz. The composition. State Graph dependency queries and outage cascades · municipal and population impact views · scouts and sentinels for signals · a cartographer in the crisis-staff meeting producing a cited decision log · a herald for briefings. What composition adds. One view linking a failing system, its first- and second-order dependencies, the responsible authorities, the latest signals and the staff’s own decisions — with provenance. What it needs from outside. The operator → system data only the state can supply, and a mandate from the crisis-management community.
5.2 A public evidence service for legislators
The need. Legislators and their staff face questions on hundreds of topics with a handful of advisers. The apps: the ENSI research engine, Hyperthesis, Scaffold. The composition. Executable research over the Evidence Graph · claim auditing · a cited-claims store · a librarian and tutor for questions · verification that blocks publication. What composition adds. A standing service that answers a committee’s question with a verified evidence pack — every figure checked against its source, every gap declared — and feeds the answer back into the graph so the next question starts further ahead. What it needs from outside. A partner office and published standards for what the service may and may not say.
5.3 A national priority register
The need. No country keeps a transparent, re-weightable register of where talent and money would do most good. The apps: Hard Frontier, Future of Politics, Climate Lab, Economy Simulator, Objective Planner. The composition. Impact-per-effort scoring by independent raters · registers of institutionally unowned issue domains · technology scoring with live re-weighting · models of how value is actually created · tests of whether a task is meaningful for a goal. What composition adds. One register in which anyone can see the evidence behind each priority, apply their own weights and watch the ranking move — with scores calibrated over time by the warden against expert panels and outcomes. What it needs from outside. Calibration against human experts, and an institutional home trusted across political camps.
5.4 A civic accountability desk
The need. Fact-checking and democratic-health monitoring are slow, underfunded and easily dismissed as partisan. The apps: Lhari, DemocracyWatch, Bohemie, Noeverse. The composition. Bounded verification of public claims · a multi-dimension democratic-health model on real data · measurement of what goes right as well as wrong · claim graphs of long-form political media. What composition adds. A desk that verifies at machine speed but publishes at human pace: claims linked to the responsible authority, verdicts proposed by the verifier and published by the steward only after an editor’s signature and a right of reply, corrections as prominent as findings. What it needs from outside. Real data pipelines, editorial staff, and cross-camp legitimacy — the hardest part.
5.5 A builders’ formation pipeline
The need. The stack depends on people who can direct agent fleets, judge their output and govern them — the skill that determines who gains from AI. The apps: Aether, Machine Game, the Agent-Driven Meetup, Talent Test, the apprenticeship idea. The composition. Courses with AI-curated reading spines · simulations of governing an AI-run organisation · live critique formats · cognitive mapping · apprenticeships inside real organisations. What composition adds. A pathway that turns motivated people into builders and stewards of the stack — tilting the balance towards broad gains rather than concentration. What it needs from outside. Accreditation and a funding model for fellows.
6 · The operating model — a studio with two sides
A civic stack of this breadth would normally be built by a public agency with a large budget or a foundation-funded institute with a large staff. The executor era makes a third form possible: a studio with a commercial side and a public side, sharing one substrate.
The commercial side field-tests and pays. The most mature agentic engineering happens where a customer can say no — governed sales workforces, research engines, context-sharing agents, meaning-level media tools. That is where the substrate’s components are invented, stress-tested and debugged, and where revenue can come from. Shared infrastructure, shared playbooks and shared agents are why a studio can outrun separate startups — and agentic tooling multiplies that logic when the shared “talent” includes agent fleets that move between ventures.
The public side opens and deploys. A public-interest organisation takes the proven components, runs them under public rules and opens them: the substrate as open-source infrastructure, the graphs as public data with provenance, the skills as published procedures, the agent roster and its “never” rules as a public document.
The flow between them is deliberate. Every civic module should be able to say which commercial product field-tested its components; every commercial product should be able to say which public module its patterns now serve. The Cybernetic Teammate result (one person with AI matching a two-person team) and The Headless Firm‘s argument that coordination costs set firm size point the same way: the minimum viable team is small — but only if infrastructure is shared rather than rebuilt.
Design for the broad-gains branch. The “regime fork” is the warning that matters most: the same coordination-compressing technology yields broad gains or superstar concentration depending on who controls the reorganised coordination. A stack holding a country’s state, evidence and discourse graphs is exactly such a structure. The counter-design is structural: public data, open procedures, published rosters, independent audit, institutional partners with their own copies, and a governance body independent of the builders.
Sustainability. Three revenue lines can carry the public side without capturing it: licences of substrate components to firms; commissioned evidence for institutions, published by default; and grants paid against measured public-value outcomes.
7 · Governance — the constitution of the stack
Civic modules must be the most governed software in the stack. The rule set is short:
A published roster — every agent’s role, tools, limits and “never” rules, readable by anyone.
Gated decisions — only the Steward publishes, only with a named human’s signature; the Warden can suspend any agent.
Visible activity — every action carries the agent’s identity and lands in a log that is publicly summarised.
Safe defaults — read-only unless granted; bounded budgets; a kill switch per agent and for the whole stack.
Measured honesty — citation accuracy, correction rates and calibration against human judgement published regularly; incidents written up as standing rules.
Built for European law from day one — documentation, human oversight and transparency native to the design, not retrofitted.
A charter drafted like a good charter — articles with worked examples and strongest objections, adversarially tested against real cases until conflicts surface, and published only when verified — the way the Sentience Accord was built.
8 · Twelve months — the build sequence
A sequence for a small team with existing executor products. It favours extraction over invention and partnership over permission.
Quarter 1 — Extract the substrate
Extract the six substrate components into one shared, versioned package.
Put every civic module under version control.
Publish the agent roster and “never” rules, drafted and adversarially tested.
Remove, or unmistakably label, every simulated or illustrative number in public-facing prototypes.
Quarter 2 — Stand up the three graphs
State Graph: refresh the state’s open registers on a schedule; request the missing operator data through a partner community.
Evidence Graph: run claim and causal-edge extraction on existing document libraries, with the Adversary refuting every edge.
Discourse Graph: give claim maps a shared claim identity; build bounded verification with a Steward gate and a right of reply.
Expose all three through the Librarian as governed protocol tools.
Quarter 3 — Ship two composite capabilities with partners
The crisis common operational picture with a crisis-management community.
The public evidence service with one legislative committee or research office.
Stand up the Warden: calibrate the Verifier and Forecaster against human judgement in both pilots.
Quarter 4 — Publish, measure, decide
Publish the audit: agent actions, approvals, corrections, citation validity, costs.
Measure public value with a value ledger and a forecast and evaluative social-return analysis per capability.
Decide which of the remaining capabilities to build next — by measured value, not enthusiasm.
9 · Risks — how the whole could be worse than the parts
Fabrication at scale. An executable research and verification process scales error as easily as insight. Mitigation: the validator library and evidence ledger are non-negotiable; the Adversary is never switched off; figures are checked against source text; blocked sources are declared.
Concentration and capture. A single stack holding a country’s state, evidence and discourse graphs is a power centre. Mitigation: public data and procedures, partner institutions with their own copies, an independent governance board, and the broad-gains design choices in §6.
Legitimacy. An accountability desk seen as partisan is worse than none. Mitigation: human editors own every verdict; right of reply; corrections as prominent as findings; good news measured as rigorously as bad; cross-camp advisory membership.
Single points of failure. A stack built by a small team depends on a few people, one cloud account and a handful of model providers. Mitigation: complete runbooks, shared ownership, multi-provider routing, and European hosting for civic data.
Operational fragility. Agent-operated deployments fail in characteristic ways — the wrong artefact shipped, a truncated build that silently deletes content, a stale snapshot reverting a fix. Mitigation: publish-blocking checks and “render and look before you report done” as standing rules for every civic deployment.
Liability. Someone must answer when a civic agent is wrong. Mitigation: the Steward gate means a named human always approved what was published; the ledger shows what they saw when they approved it.
Close
Cheap building lets a very small team make many things. The greater story is what those things become when they stop being separate. The parts of a civic operating system — a way to see the state, to know what is true, to make sense of public argument, to decide what matters, to hold power to account, to govern the agents doing the work, and to form the people who will run it — are already patterns that AI-native builders know how to make. What they lack is the shared substrate, the connected graphs, the public rules and the institutional partners that turn a collection of products into infrastructure.
That is a job for a public-interest organisation, not a startup: to take the patterns the market has proved, open them, govern them more strictly than any commercial product, and deploy them with the institutions that already want them. Eight Plays for the Civic Apps sets out which pieces to carry first, and Civic Apps: Value for Society how to prove, in terms a funder or a ministry will accept, that they created public value.



