The Policy Innovation Lab believes that AI, used responsibly, can improve policymaking, governance and service delivery. Members of the Lab recently published the CRAFT principles for responsible AI use in government. In this article, we demonstrate an example of how the Lab put these principles into practice when using Large Language Models (LLMs) to analyse South Africa’s national budget. We begin by looking at the analysis itself before discussing the methods for responsible LLM use.
Analysing the National Budget
The national budget is the starting point for resourcing the South African government’s plans, policies and services. No matter how good a policy is, if the national budget does not allocate resources towards it, then it will not be implemented. It is therefore essential that the budget allocates its limited resources in ways that align with government’s medium and long-term priorities.
The Lab was recently asked to support The Presidency of South Africa to assess the alignment of the 2026 executive budget with the three priorities laid out in the Medium-Term Development Plan (MTDP), which we will call P1, P2 and P3. These are:
P1: Inclusive economic growth and job creation,
P2: Poverty reduction and the cost of living, and
P3: A capable, ethical and developmental state.
The analysis consisted of both quantitative and qualitative elements. The main approach was to go through each budget line and use an LLM (in this case, ChatGPT 5.6 Sol Pro) to label these according to which priority that budget item was most aligned to and according to how that budget contributed to those priorities according to the following definitions:
Category
Meaning
Examples
Direct
Produces an identifiable service, benefit or output for people, firms, communities or public institutions aligned with P1, P2 or P3.
Paying social grants; providing school meals; constructing houses; issuing identity documents; supplying water to communities; funding business support directly received by firms.
Enabling
Provides the infrastructure, people, equipment or systems needed to deliver a priority result.
Training frontline officials; procuring medical or laboratory equipment; maintaining an information system used to process grants; providing vehicles for service delivery; upgrading facilities used to deliver public services.
Institutional
Improves the policy, regulatory, planning, monitoring, research or oversight environment.
Developing regulations; conducting policy research; monitoring programme performance; producing official statistics; evaluating service-delivery outcomes; undertaking inspections or regulatory oversight.
Operational
Keeps the department operating but is not readily attributable to a specific MTDP result.
General human-resource management; internal finance and supply-chain administration; office accommodation; routine legal services; internal communications; executive support not linked to a specific priority result.
Not Aligned
The link to a priority is too weak, indirect or insufficiently documented.
A project with no stated MTDP outcome or target; a general awareness campaign without a defined priority result; membership fees with no documented contribution to P1, P2 or P3; legacy activities whose current policy purpose is unclear.
The model was provided with speeches and Estimates of National Expenditure to help it determine what each budget item was for. This was repeated for historical data (going back to 2022) and for future budget projections to identify historical and emerging trends in budget allocation.
The headline results of the analysis can be summarized in the Figures below.
Figure 1: Total budget allocation to each priority, showing what proportion of expenditure is Direct, Enabling, Institutional, Operational or Not Aligned.
Figure 1 shows that most of the national budget goes towards priority 2, poverty and the cost of living. Indeed, the vote with by far the largest single budget allocation (R302.4 billion) is towards Social Development, including social grants. This money directly addresses poverty. This is also seen in Figure 2, where each vote is broken down into the different types of expenditure.
Figure 2: Contribution pathway profile for each executive Vote. Each horizontal line shows how the Vote is divided between Direct, Enabling, Institutional, Operational and Not Aligned expenditure. Votes are ordered from the highest to the lowest share of expenditure classified as Direct.
Using LLMs in alignment with the CRAFT principles
While a lot more can be said about the budget and the subsequent analysis, the focus of this article is to ask the question of whether the analysis should be trusted, given that it relies on an LLM to label and categorise budget items, a task it was not specifically trained for. Indeed, studies show that LLMs may differ in how they label and categorise qualitative data (called coding) compared to human experts and may classify the same text differently on repeated runs, even when the temperature (parameter used to control the degree of randomness in the outputs of an LLM) is set to zero [1] [2][3].
Despite these risks, the gains in efficiency, and sometimes quality, mean that LLMs are increasingly used to analyze public-sector qualitative data. The Lab recently published the CRAFT principles for using LLMs for this type of analysis[4]. These principles are
Controllability – Rigour – Accountability – Fairness – Transparency
and guide the responsible use of LLMs in the public sector (you can find the CRAFT principles here or read our discussion here). Here, we discuss how we translated each of these principles in the budget analysis.
Controllability
The first principle recognises that people and not machines must remain in control. Civil servants should retain the ability to understand, guide and perform the task independently of the LLM.
For the budget analysis, the people conducting the analysis defined the policy question, the MTDP priorities and the five contribution categories before the model was used. The LLM applied a human-designed coding framework to existing budget lines. Because each classification remained linked to its original budget item and source material, analysts could inspect, challenge and revise the model’s judgement rather than accept it automatically. The analysis was also not once-off, but a trial run was done. The user made observations and identified improvements to the methodology, output format and source data, thus remaining in control of what the LLM was and was not trying to do.
Rigour
Since LLMs may hallucinate, drift or be sycophantic, every AI-generated claim should be verified against reliable evidence, trusted sources and subject-matter expertise before it influences policy decisions.
Rigour was supported by grounding the coding in official South African evidence. The model received the relevant Estimates of National Expenditure and departmental budget speeches, rather than being asked to infer the purpose of a budget item from its title alone. The same definitions and coding process were then applied at line-item level to data going back to 2022 and to the forward estimates before the results were aggregated. This created a consistent basis for comparison and made unusual classifications easier to trace back to the underlying evidence.
The LLM was instructed to provide estimates for how confident it was that a label was correct and also flag low-confidence items for human review (there were 169 items flagged, about 4% of labels). These items were reviewed, and it was found that the budget items were ambiguous in nature, making it difficult for both the LLM and the human to accurately apply labels, so only one of them was changed in review. Furthermore, quantitative analysis required the LLM to write, execute and save Python code. This code was then reviewed and, sometimes, re-run, using intermediate datasets also created by the LLM during its analysis to ensure provenance. Spot checks were done to verify calculations and totals were compared with the values from the Appropriation Bill and other National Treasury resources.
It is worth mentioning that the nature of the budget vote categorisation can be subjective and ambiguous; expenditure one person thinks directly addresses a priority (according to our definition), another person may think only enables it, or that it doesn’t align with the priorities at all. Through iterative testing, the category definitions could be improved; although it seems unlikely that there would be no instances of ambiguity or overlap. This further highlights the need for rigorous human oversight.
Accountability
The user of an LLM should remain accountable for decisions, especially in democratic governments. Even where AI contributes to drafting or analysis, a clearly identified person must remain responsible for the final advice, recommendation or decision.
The LLM produced classifications, not decisions. The Lab remained responsible for designing the method, selecting the evidence, interpreting the results and communicating the limitations. The analysis was used to support policymakers’ judgement about alignment with the MTDP, and the user ensured that the scope did not creep to other applications, so that they could take full accountability for what was produced. Keeping the analytical and decision-making roles separate ensures that named officials and researchers, rather than the model, remain answerable for conclusions drawn from the work.
Fairness
AI systems reflect the data on which they are trained, which tends to be disproportionately Western in both quantity and quality. Thus the LLMs may leave out perspectives or concepts that are more suitable in South Africa’s context.
To reduce the risk of importing inappropriate assumptions, the model was grounded in South African policy categories and locally produced government documents. The MTDP priorities, departmental speeches and Estimates of National Expenditure provided the relevant institutional and social context. The ‘Not Aligned’ category was also important: it allowed the model to acknowledge that the evidence was insufficient instead of forcing every item into a priority or categorical label. This does not eliminate bias, but it limits reliance on generic, predominantly Western concepts and makes context-sensitive disagreements visible for human review.
Transparency
Citizens deserve to know when and how AI has informed government decisions. Transparency does not require disclosing every interaction with an AI tool, but it does mean explaining its role in proportion to its influence on a decision. Transparency builds trust and enables meaningful public accountability.
Transparency was built into the analysis by documenting the model used, the task it performed, the definitions applied, the evidence supplied and the years covered. The prompt was loaded as a document as opposed to in the chat, allowing better management and forcing more reflection on the purpose and iterative changes. The published figure reports aggregated results, while the underlying method retains a line from each total back to the individual budget items and their classifications. The article also states the known limitations of LLM coding, including disagreement with human coders and variation across repeated runs. Readers can therefore distinguish between official budget data, human methodological choices and AI-assisted classifications. All source data, prompt instructions, intermediate datasets and Python code were made available to the recipients of the analysis.
By adopting the CRAFT principles, the Lab believes that LLMs can and should be used responsibly to improve the efficiency and effectiveness of South African public sector qualitative analysis and policymaking more broadly.
[1] Parfenova, A., Marfurt, A., Pfeffer, J. and Denzler, A. (2025). Text Annotation via Inductive Coding: Comparing Human Experts to LLMs in Qualitative Data Analysis. ArXiv, abs/2512.00046. doi: 10.18653/v1/2025.findings-naacl.361.
[2] Nguyen-Trung, K. and Friese, S. (2026) ‘Towards a methodologically congruent framework for GenAI use in nonpositivist qualitative research’, SSRN Electronic Journal. doi: 10.2139/ssrn.5874482.
[3] Coqueret, G., Llull, J., Oswald, F., Pérignon, C., Scheuch, C. and Vilhuber, L. (2026) ‘Randomness in large language models: What researchers need to know (and report)’, arXiv, arXiv:2607.24372.
[4] Fourie, W., Manicom, G. and de Villiers-Botha, T. (2026) ‘The CRAFT principles for the responsible use of large language models in policymaking’, arXiv, arXiv:2607.15704. doi: 10.48550/arXiv.2607.15704.
