Strong evidence for policy decisions
We are often reminded to pay attention to the evidence, but don't think about what good evidence looks like

My last article may have left some readers a little puzzled. I argued that evidence is important for decisions but evidence-based policy doesn’t work as intended. This leaves some gaps that need explaining. To start with, if evidence isn’t enough, what should we base our policy decisions on? And, more philosophically, what does good evidence actually look like?
There are a range of views about these questions in the broader world of public policy, but, in my experience, not many of them are consistent with practice. Even fewer of them have a solid basis in epistemology, or the logical structure of human knowledge. For this reason, my aim here is to think things through from first principles, and draw inspiration from other fields with strongly stress-tested knowledge practices.
The components of policy making
The previous article touched lightly on the question of what we base policy decisions on. My views line up neatly with a simple summary provided recently by Sean Innis. He has written a series of articles on the ‘Policy Recipe’, where he argued that policy decisions are made from three ingredients: values, pragmatism and evidence. While I’m not convinced by the metaphor, his three components make sense.
As I noted in my previous article, values are essential to policy decisions (which was why evidence alone is not sufficient). A policy is always designed to make the world better in some way, and the definition of ‘better’ always depends on some set of values. These values may be largely self-evident (say, it is better to live than to die) but there are always groups of people who espouse the opposite value. The challenge is that there is no way of deciding these objectively.
Pragmatism is essential. A policy that is implemented is better than a brilliant policy idea that never goes anywhere. It focuses on the question of what can we actually achieve at this point in time, given all the various constraints. This can include everything from laws, to budgets, to public opinion, to constitutional structures, to the capabilities of people involved.
Evidence covers what we have learnt from experience, what we know about what works and what we should expect to happen.
Sean used the analogy of a recipe, but an alternative is to think of policy as akin to an architectural drawing.1 The architect is always guided by values, whether it is an aesthetic, intentions for use, goals, and similar considerations. They then have to work within all the pragmatic constraints: site, location, weather, budget, time, etc. And it is critical they take into account all the evidence about what does or doesn’t work regarding engineering, structural features, sun and weather and so on. All three are core considerations, and, as an aside, there is a fourth that our modern world adds: all the laws and regulations that need to be complied with. This is equally relevant for policy decisions.
What is policy evidence?
What counts as evidence for a building is fairly clear. We mostly want to know if the building will stand, not have issues and be able to be built for the cost we can afford. Evidence for policy is less clear cut. What counts as valid evidence? On what basis should we support a decision that is pragmatically possible and lines up with our agreed values? How do we know if it is going to work or not?
As noted in the previous article, it is common to cite systematic reviews, Randomised Control Trials (RCTs) or other forms of rigorous academic work as the gold standard for evidence. However, in an important way, this focus on facts and data misses the point of what we need evidence for. It might be scientific in some senses but it typically doesn’t give us the same kinds of products as truly useful science.
To see how this works, it helps to go back to first principles and look at the logic underpinning what we need evidence for when we are making a policy decision. At an abstract level, any policy is designed so that certain people take particular actions to change what is happening for the better. A good policy has to, at a minimum, get the logic that connects all of this together correct. That is, when we put the policy in place in the real world, the relevant people do the sorts of things that lead to the types of positive changes we anticipated. It is similar to building the house. We want evidence to show that the house will actually stand and look like the drawings.
Values tell us what we want those positive changes to be, and pragmatism helps us know we can actually do. But we need the evidence to tell us if our planned policy will actually lead to the changes we want.
So, ideally, evidence will provide us a reliable way of identifying what people (or systems or organisations) will do in response to various policy actions (or the status quo). That is, it should tell us what will happen if we implement policy A versus policy B and so on. It follows that the perfect evidence would be some kind of fully reliable prediction machine or crystal ball to tell us what would happen in various situations. However, such a machine doesn’t exist for us mere mortals (despite what some proponents of generative AI might suggest).
This isn’t a new problem or one specific to policy. We deal with it in all fields and there is a standard scientific approach. What science typically does is work hard to build broadly applicable and reliable theories, or (in some cases) models, that we can use to predict what we expect to happen in the relevant specific circumstance. Strong theories and models allow us to plug in a range of different situations and see what the likely real world effects or outcomes are. For example, the scientific theories that tell us how diseases are spread by bacteria or viruses through various fluids have allowed many improvements to health care and human longevity. We can evaluate the likely effects of different interventions without needing to trial every one of them individually.
Importantly, the benefits of scientific theories or models go beyond their predictive abilities - and beyond what a crystal ball might give us. They often provide us with an understanding of why or how certain things will happen; they explain the mechanisms or dynamics that lead to particular outcomes and not just the outcome itself. The understanding allows us to identify ways of intervening positively or pick where to leave things alone. Down the track, understanding the mechanisms or dynamics also allows us to diagnose why something went wrong. In the absence of this, we can only guess, try and hope.
The logical structure of a policy decision, therefore, requires evidence in the form of something that explains what is likely to happen and why it will play out that way. This allows us to assess different options and anticipate real world changes when we implement something.2
Good evidence should be scientific, in the proper sense of the term
The ideal evidence for a policy decision is, therefore, a robust scientific theory (or perhaps model) that is reliable and directly applicable to the policy issue at hand. In practice, this is rare and we often have to use multiple overlapping theories of various grades of reliability. Importantly, these are logically distinct from the common outputs of systematic reviews or RCTs. Those outputs typically tell us what tends to happen in past situations, within a statistical standard of reliability.3 However, they less commonly tell us anything about why these results hold or the mechanisms at play.
In strong scientific practice, facts, data, RCTs and reviews help us decide between competing theories and identify which are the most reliable or most likely to be true. If our best theory is contradicted by the data, we need to find a better one. Where we don’t have a compelling or strong theory, we may have to rely on the experimental evidence as our guide but this has intrinsic limitations. It tells us what happened in a particular situation in the past. Without the reliable theory, we can’t be sure that it will happen in the same way in our current context as we don’t know why it happened that way.
In the world of public policy, it is common to hear complaints that decision makers, especially politicians, rely on anecdotes they have heard rather than the data that is available. An anecdote is only one data point and therefore data should always outweigh an anecdote. What these complaints miss is that anecdotes match the logical structure of human decision making better than data.
An anecdote is always a data point plus a story. The story provides a critical role as it provides an account of why something happened. The why is logically important for decision makers as it allows them to naturally play forward different scenarios and test them against policy options. Data, on the other hand, might provide us with robust correlations but it is much harder for humans to think it through for different scenarios. Data doesn’t provide the why. It gives us no understandable account of how the world is and what will change.
Strong policy evidence, therefore, should be a rigorously justified theory that explains how the particular part of the world works and what the effects of policy changes will be. Everything we normally talk about as evidence is important here, it should just justify our theory rather than be all there is to evidence. When evidence is logically structured as a theory, explanation or account of the world it is much easier for humans to use it to make good decisions. Facts, data and charts are harder as they don’t fill in the gaps of how or why something happens.
There are other benefits to having evidence in the form of a theory or explanation, rather than statistics and data. One notable one is that it provides a conceptual blueprint for how the policy will play out in practice. That is, it provides us with the ‘theory of change’ for how the policy will succeed to achieve the desired positive outcomes. Having this clear makes it easier to change course if things aren’t working and empowers us to evaluate outcomes more simply at the end. Once we understand how we think something will roll out, we can check to see if we are correct.
To summarise, we need to downplay the role of facts and data in our policy making so that we can make policy evidence more scientific. Science is ultimately about providing theories and models that explain the world, not data and charts about what happened. Of course, science doesn’t land on these theories without significant effort and uncertainty, and some fields never achieve certainty. A more scientific approach to evidence for policy decisions includes greater comfort with uncertainties and unknowns.
This analogy shouldn’t be taken too far. There are real risks with using building or mechanical analogies for policy work as they can lead us to think we can change human behaviour in the same way we can shift a wall in a house. Humans have will and agency. Walls don’t.
In a related point, there are interesting arguments that we also need to explain the why when we are capturing processes and standard operating procedures. If we just describe the what, it is often hard to achieve the same outcomes. See this Substack article for more details:
There are growing questions about the usefulness of statistics for scientific work, as I wrote about here:
Lies, Damn Lies and Probabilities
Recently, Sean offered an interesting hypothesis in response to a post. He wondered whether “The thinking limit of an AI reflects the limit of statistics as a way of seeing the world.” I’m not sure if the idea is right, but it seems like a significant insight. For there are a range of reasons why statistics are inher…



Loved this.
Your definition of 'strong' evidence feels powerful, at least in the policy realm. I have a quibble though. I would define what you describe as a strong policy model rather than strong policy evidence. Evidence, for me anyway, cannot make a temporal leap from past to future without changing into something else. It is why in the piece you mentioned (thanks by the way) I ultimately shifted from evidence to understanding.
I also loved the architecture analogy. I think it works well. But I might stick with the recipe, partly mainly the laws of physics are less obviously defining of the outcome!
Lots to talk about next time we see each other.