
You’ve probably heard of the government’s own AI tools. Beyond them is a whole field of others — from universities, non-profits and companies alike — that policy makers can use now. The trick is finding them, and telling the good from the extractive.
If you work anywhere near government, you’ve likely heard of Humphrey — the government’s own bundle of AI tools for civil servants, which includes tools like Parlex, for searching decades of parliamentary debate in seconds, and Minute, which formalises the output of official meetings. The most famous of them was Redbox, a chatbot for drafting and summarising — which the government has already quietly retired, as enterprise tools like Microsoft Copilot spread across Whitehall. The rest are valuable and safe to reach for, having been built by government for government.
But Whitehall’s own kit is only part of the story — and, as Redbox shows, not always a permanent one. Beyond it sits a much wider field of AI tools: some built by universities and non-profits for the public good, some by companies with a product to sell. Any of them could be put to work today, provided you understand what each one can and cannot do, and can tell one from the other.
The sections below follow the arc of making a policy: spotting a problem, making sense of the evidence, designing a response, winning consent, and knowing whether it worked. For each, I’ll point to tools worth knowing and be clear about where the health warnings are.
Spotting the issues new policies need to address
Good policy makers are always scanning the horizon, trying to spot emerging problems before they become crises. There’s serious craft to this already: the Government Office for Science’s Futures Toolkit gives officials a dozen structured methods for thinking about what might be coming.
The difficulty isn’t method, though — it’s speed. Government tends to discover problems late: through a media storm, a select committee hearing, or a statistic published months after the fact. The people living the problem knew about it long before Whitehall did. That is, at root, a speed-of-information problem, and it’s exactly the kind of thing AI can help with.
Two ways in. The first is forecasting: estimating how likely a future event is, and putting a number on it. On platforms like Metaculus, AI agents are getting better at this: last September a British startup, ManticAI, became the first bot to finish in the top ten of an international forecasting competition, beating most of the humans.
The second is hearing what people are worried about before it reaches official channels. Social listening uses AI to analyse what large numbers of people are already saying in public. Most commercial social listening tools are built for brand managers, but public-sector-facing versions exist too. UN Global Pulse advocates for the use of social listening to support service design; in one example it uses AI to transcribe rural radio phone-ins so local concerns can be mapped to inform aid provision. In the UK, Orlo (a for-profit) is used by councils, NHS trusts and police forces to understand community sentiment and support service design.
What to watch out for, and using it safely:
Always check the motive behind a prediction engine. The real money in forecasting is in betting (Polymarket, Kalshi and the rest), and a system built to win bets has different incentives from one built to inform a minister.
The best results still come from keeping a person in the loop — particularly where the data is partial or messy. One of Metaculus’s top forecasters calls this a “centaur” approach: humans frame the question; AI helps forecast the answer. As he puts it, “asking the right questions is an art that AI has yet to master”.
Making sense of the evidence
Solid policy rests on good evidence about what actually works. But reading the research is slow, specialist work, and traditional reviews often arrive after the field has already moved on.
This is the problem living evidence is meant to solve — a summary of the research that updates itself continuously as new studies appear, instead of being rewritten from scratch every few years. Cochrane, the organisation behind medicine’s most trusted evidence reviews, is doing serious work to make living evidence the norm, and AI makes continuous updating practical at scale. Tools like Elicit speed up the searching and sifting that used to eat weeks. And the idea is already live in the UK: the Health Equity Evidence Centre uses machine learning to keep “living evidence maps” of what works to reduce health inequalities in primary care.
Some emerging tools go further — Consensus’s “Consensus Meter” scores how much the top studies on a yes/no question point the same way. It’s a fast read on the weight of evidence, but it weighs only the top 20 papers, can misclassify a finding, and can miss the nuance in your question (a study about children counted as a “yes” to a question about adults) — so still some way to go before it’s government-ready.
What to watch out for, and using it safely:
Speed is not the same as trust. On something as basic as judging whether a study is biased, AI and human reviewers currently agree only somewhere between four and seven times out of ten. Use the tools to do the legwork, then verify, and never let the machine’s confidence stand in for your own. This is exactly the territory Cochrane’s RAISE guidance was written for — Responsible AI use in Evidence Synthesis, a practical standard for deciding when and how to use these tools, and when not to. We’ll come back to its central rule at the end.
Designing and testing new policy interventions
Once you know the problem, you have to design the response — and, ideally, work out who it helps and who it hurts before you commit. Today that’s either done by rule of thumb or by slow, specialist economic modelling.
This is where the standout public-benefit tool lives. PolicyEngine uses microsimulation: building a virtual population of thousands of representative households and running a proposed tax or benefit change through it, so you can see who gains and who loses before you commit. It’s open-source, non-profit, and good enough that the data science team inside Number 10 has built on it.
For a different kind of question — how a whole system behaves, not how a change lands on individual households — there’s systems modelling: simulating the feedback loops of an entire system to test “what if” at the level of the whole. En-ROADS, built by the non-profit Climate Interactive with MIT Sloan, lets you pull the policy levers on climate — a carbon price here, a shift to renewables there — and watch the modelled effect on emissions, prices and temperature ripple through. It’s free, and its negotiation-grade sibling C-ROADS has been used in real international climate talks.
What to watch out for, and using it safely:
Beware the tempting shortcut of “silicon sampling”, where you ask a large language model to role-play the public and tell you what they would think. It flatters you with plausible answers while quietly stereotyping the very groups it is pretending to be. Fine for sharpening a survey; no substitute for asking real people.
Iterating design and winning consent
No policy survives contact with the public if there isn’t broad support. Consultations and citizens’ assemblies are the main ways that is tested now. But citizens’ assemblies are expensive and slow, so they are used rarely; and analysing thousands of consultation responses by hand can take weeks while still over-weighting the loudest voices.
Government already has a foot in this door: Consult, one of Humphrey’s tools, reads through the thousands of responses a public consultation attracts and clusters them into themes for officials — the sort of job that used to swallow weeks of analyst time.
Beyond that, three tools are emerging that can help with engagement, deliberation and decision-making at scale. Used well, they don’t just test a finished proposal; they feed back into the design, making consultation part of the iteration loop rather than a one-off gate at the end.
Talk to the City (run by a non-profit) helps with engagement: it reads a large pile of free-text responses and clusters them into themes, showing not just what people think but how their views group together — so you can hear thousands of people at once, a hearing aid rather than a mediator. In 2025, the peace foundation CMI used it in Yemen to gather young people’s views on political participation and peacebuilding across 18 governorates, safely and in Yemeni Arabic, with a 94% response rate.
Polis (also non-profit) supports deliberation: people write short statements and everyone votes agree or disagree, and the statistics reveal both the opinion camps and, more usefully, the statements that command support across all of them. It has real policy miles on it — Taiwan’s vTaiwan process used Polis to broker a public consensus on how to regulate Uber, which the transport ministry then wrote into the rules.
And Google DeepMind’s Habermas Machine reaches towards decision-making: the AI drafts candidate “common ground” statements, people critique them, it revises, and the group ranks the result — an approach that, tested on more than 5,000 people in the UK, produced more agreement than human mediators managed.
What to watch out for, and using it safely:
The real risk here is not the technology, it is consultation theatre — analysing responses beautifully and then changing nothing. A tool that helps you hear people only earns its keep if what they say actually moves the decision. And the Habermas Machine, for all its cleverness, is corporate research rather than a public-benefit tool, so it needs careful handling.
Knowing what works — the feedback loop
The last stretch is finding out whether the policy did any good — the stage where AI may be most useful and least glamorous. Formal evaluation often reports years after launch, long after the decisions it might have shaped.
Some international bodies are furthest ahead in using AI to close that gap: the World Bank has used machine learning to read hundreds of project reports at once, and an international research team recently ran a major global evaluation of 1,500 climate policies, using the OECD’s Climate Actions and Policies Measurement Framework to find the handful — just 63 — that delivered major emission cuts. However, there isn’t yet a public-benefit evaluation tool a policy maker can simply pick up, the way there is for microsimulation or social listening. This is still bespoke work.
Coming down the pipe is what evaluators call rapid-cycle (or real-time) evaluation: using signals services already generate — complaints, call patterns, missed appointments — to see within weeks whether a policy is on track. Real-time sentiment analysis does exist, but there’s work to do to ensure it’s robust enough to use at scale in public service.
What to watch out for, and using it safely:
Real-time signals can tell you something has changed, not why. A spike in drop-offs might mean the policy is wrong, or it might mean the form is glitching. Both need fixing, but only a human can work out whether the problem is technical, operational or substantive. And some impacts simply take time to show: prevention programmes, school reforms and behaviour change can take years to prove their worth. Rapid-cycle methods tell you whether something may be off track; they do not prove whether the policy ultimately worked.
Keeping AI in the right place
Run back through those examples and a single rule connects them: these tools are there to help you see, hear and understand — not to decide. Hold onto that, and a short checklist keeps you on the right side of the line.
1. Read the label. Four questions separate the trustworthy from the dubious: who funds it; whether it trains on your data; whether it explains how it works and handles bias; and who is accountable when it gets something wrong.
2. Keep a human holding the pen. The good tools are copilots, not autopilots. Responsibility for the decision stays with the person, and the workflow should make that unavoidable.
3. Use it to see, not to settle. AI is good at finding patterns and bad at knowing what they mean. Let it tell you complaints are rising; do not let it decide whether that is a scandal or a form-design glitch. The framing, and the judgement, stay yours.
The tools worth searching for
Not one of these tools will make policy for you. That is not a flaw to be fixed; it is the design. The good ones leave judgement where it belongs: with the policy maker. And whoever built them — a lab, a charity, a company — the ones worth having are rarely the ones with a sales team behind them. No launch, no pitch with your name on it; you have to go looking.
That is rather the point. The tools worth knowing about are usually the ones you have to go looking for. Knowing they exist is half of it; knowing how to read the label is the other half, and it is the sort of question I spend a good deal of my time helping government work through.
A note on the tools named here: these are examples, not endorsements. I’ve pointed to each one to show what an approach looks like in practice, not to recommend it for any particular job, and I have no stake in, or commercial relationship with, any of them. This is a fast-moving field — tools improve, change hands, and sometimes vanish altogether — so treat every name as a prompt for your own enquiry, not a shortcut past it, and run it through the checklist above before you rely on it.
