Civil servants are increasingly using language models to read and analyse policy documents. Dr Aleksei Turobov and Prof Dame Diane Coyle tested four ways of doing this on a corpus of 63 United Nations texts, and found no technical criterion for which method to choose — the decision has to be made by humans, not the software.

Officials who grant a language model access to their documents are choosing among options: an AI chat window into which an analyst pastes a consultation response, a script that pushes an entire archive through a commercial AI model, or a statistical model that sketches a map of a corpus nobody has time to read. In our research – published in Data and Policy – we ran a body of 63 documents through four modes of use. We find that even with a fully automated approach, humans still need to interpret results and decide the outcome.
The choice of how a model reads and analyses is a governance decision that should weigh trade-offs between speed, cost, depth and oversight. That judgement shapes the evidence a government will act on, and this accountability is why it cannot be left to whoever signs the software contract.
Using a corpus of 63 English-language United Nations (UN) policy documents on artificial intelligence (2017 – 2024), we focus on thematic analysis as a standard procedure in which an analyst labels meaningful passages and so groups them into themes. This was our ‘universal’ policy analysis use case, which we tested across: (a) manual expert thematic analysis (five experts); (b) a statistical topic model that scanned the whole corpus for clusters of words that travel together; (c) a custom version of GPT-4, working from a written protocol we first set out in a working paper yearlier, and; (d) python script of using GPT API, sending every document to a model and returned codes, quotations and themes.
The human experts set the standard for depth, so we could ask what the other approaches added. We found that the experts caught what only close reading catches: the relationships a text leaves implicit and the assumptions it takes for granted. AI models missed important context.
The assisted workflow, using custom ChatGPT with a specific prompt, most successfully balanced depth with scope. The model returned 713 distinct codes across all 63 documents, each anchored to a quotation that could be checked. We then built 12 themes from them, from the UN’s role in AI governance to geopolitical tension. In this case expert effort shifted from extraction to interpretation, and none of it required programming. At times ChatGPT produced artefacts, meaning spurious output; it followed instructions inconsistently; the 2024 model we used hit token limits on very long documents; the chat interface gave us no control over its settings; and it needed manual validation throughout.
The most counter-intuitive result came from the scripted pipeline. It fed each document to the model with fixed instructions, generated 473 precise codes with quotations, and wrote them to a spreadsheet without manual control over outcome. Because it analysed every document on its own, its raw output contained 57 overlapping themes. Reducing those to the nine we report took two further passes, one merging themes and one resolving overlaps; we set the consolidation criteria and validated the survivors against the underlying codes and quotations. Whoever writes those criteria decides which themes exist – that is the authorship. A ‘fully automated’ system still needs human governance; automation moves that layer to the set-up of the pipeline, which we tuned by hand, and to the end.
That is why we treat our written protocol, reproduced in the article’s appendix, as a governance document. It specifies the model’s role, the steps to follow, the output format and the rule that every code carries a quotation. For any institution using AI, it is imperative to establish consistent protocols and maintain records to validate how that evidence was produced.
The topic model sits at the other pole, producing 14 rough topics at minimal cost and with minimal oversight. It supplies no quotations to support its claims, so it is at best a tool for macro-level awareness to navigate and inform further focus.
While all four workflows covered the same core areas – human rights, ethics, security, international cooperation and sustainable development – they differed in their trade-offs between speed vs depth and control over interpretation.
However, both AI-model-based workflows sent documents to a third party’s commercial model. Where an institution’s data rules restrict that, those two workflows that combined scale with quotations would be ruled out. We view commercial application programme interfaces (APIs) as a transitional stage leading toward state-owned analytical capabilities – such as open-source or custom models running on in-house infrastructure and fine-tuned on the organisation’s own documents. While this approach requires more resources, it grants the responsible authority greater control over its own analysis.
The four workflows are all divisions of labour between people and software, involving properties of the institution, whatever the model. A better AI model does not settle who writes the consolidation criteria, whose data rules apply, or who is answerable for the themes that reach a briefing. Even if the balance shifts as models improve, these trade-offs will remain.

Each department’s accountable officer should sign off on the workflow in use, with the trade-offs acknowledged and a clear rule about whether documents may leave the building. Analytical units should default to the assisted workflow: the model does the initial coding, analysts develop the themes, and interpretation stays with the people. Every analytical product built with a model should have its own explicit protocol; funding the tool without funding the people who write and test protocols involves taking on the risk without the benefit. And the centre of government should decide whether the state will host its own models; until it does, data rules and analytical needs will keep pointing in opposite directions.
Language models have made policy-related decisions faster and easier to overlook. It should be made on purpose, by the people who will answer for the evidence.
Read the full article: Turobov, A. and Coyle, D. (2026). “Governing the evidence base: an empirical framework of AI archetypes for policy analysis.”, Data & Policy (open access):
Read previous working paper: Turobov, A., Coyle, D., & Harding, V. (2024). “Using ChatGPT for thematic analysis.” arXiv preprint
The views and opinions expressed in this post are those of the author(s) and not necessarily those of the Bennett School of Public Policy.