A funder asks the question every program dreads in its plainest form: did it work? Behind it sits a stronger question the program is often too quick to answer yes to; did we cause the change? The difference between those two is where a lot of evaluation quietly loses its honesty, and where a careful evaluator earns their keep.
Most real programs run in the open air. The people served are also aging, moving, finding work, losing it, meeting a good neighbour, catching a bad month. The economy shifts, a policy changes, the season turns. Against all of that, a single program is one influence among many. To claim it caused an outcome is to claim you have ruled the rest out. Usually you have not, and usually you cannot.
Attribution asks whether the program caused the change. Most designs cannot answer that. Contribution asks whether the program helped cause it, and how much we can defend saying so. Most designs can.
01 · The question under the questionDid it work, and how would we know
"Did it work" is really two questions wearing one coat. The first is descriptive: did the outcome improve? The second is causal: was it because of us? A pre and post measure can answer the first. It cannot, on its own, answer the second, because it has no account of what would have happened anyway.
That missing account has a name: the counterfactual, the world in which the program did not run. You never observe it directly; the same people cannot both receive and not receive the program. Every causal claim in evaluation is, underneath, an argument about that unobserved world. The strength of your design is just the strength of that argument.
02 · The comparison ladderThe claim each design supports
Designs are not better or worse in the abstract; they buy you different amounts of the counterfactual, and therefore different claims. It helps to see them as rungs.
- 1After only.You measured once, at the end. You can describe a state of affairs. You cannot claim change, because you have no before; and you cannot claim cause.
- 2Before and after, one group.You can show the outcome moved. You still cannot say the program moved it, because everything else moved too over the same window.
- 3A comparison group you did not randomize.A wait-list, a matched group, a neighbouring site. Now you have a rough stand-in for the counterfactual. You can make a defensible contribution claim, as long as you take seriously how the groups differ.
- 4Randomized assignment.Chance decides who receives the program, so the groups differ only by luck at the start. This is the strongest ordinary claim to cause; it still needs its limits stated, and it can still be undone by who drops out.
Evidence: where you sit on this ladder is set by the design, not by how much the outcome moved. A huge before-and-after jump on rung two is still a rung-two claim. A modest difference on rung four is still a stronger causal statement. Confusing the size of the effect with the strength of the claim is the most common way good news gets oversold.
03 · Contribution is not a consolation prizeThe honest middle
Evaluators sometimes treat contribution as the word you reach for when you failed to prove cause. That framing is backwards. For most community and public programs, contribution is not the fallback; it is the correct and defensible claim, and reaching past it to "we caused this" is the error.
Contribution analysis, as the evaluation literature developed it, makes the reasoning explicit rather than hand-waved. You lay out the theory of how the program should produce the outcome, gather evidence for each link, name the other plausible explanations, and then show why the program remains a credible part of the story once those rivals are weighed. The claim you end with is careful and it is usable: the program plausibly contributed, here is the chain, here is what would have to be true for that to be wrong.
A defensible "we contributed" outranks an indefensible "we caused it" every time a serious reader is in the room, and the serious readers are the ones who decide.
Belief: a decision-maker can act perfectly well on a strong contribution claim. What they cannot act well on is a causal claim that collapses the first time someone asks how you ruled out the economy. Overclaiming does not make the finding stronger; it makes it fragile, and it spends trust you will want later.
04 · The claim checkerName the strongest claim your design supports
Before you write "the program improved" anything, it is worth stating out loud what your design can actually carry. The tool below does that in the small: tell it what comparison you have and how exposed you are to other explanations, and it names the strongest honest claim and hands you a sentence you can defend. It runs entirely in your browser and keeps nothing.
Claim checker
Two questions about your evaluation. The result is the strongest claim the design supports, not a verdict on the program.
05 · Saying it out loudLanguage that claims exactly what you can defend
The honesty has to survive the walk from the appendix to the executive summary, which is where claims get rounded up. A few templates keep the wording matched to the design.
None of these is longer than the overclaim it replaces. They just spend their words on precision instead of on confidence they cannot back.
06 · The counterargumentContribution can become an excuse
Taken lazily, "we contributed" is a hedge that can never be wrong, and a claim that can never be wrong is not evidence; it is comfortable noise. If every program contributes to every good thing, the word has stopped doing any work.
The discipline that keeps contribution honest is the same one that would have kept a causal claim honest: name the rival explanations out loud, and say what evidence would count against you. A real contribution claim is falsifiable. It says here is the chain we think ran, here are the other stories that could explain the same result, and here is why we still think ours is part of it. That is a claim a reader can push on. "We probably helped" is not.
Claim the most your design can carry, and not one word more. The credibility you keep is worth more than the credit you skip.
Recommendation: decide the strongest claim your design allows before you see the results, and write it into the plan. Deciding after you know the numbers is how a rung-two design ends up carrying a rung-four sentence. Set the ceiling early; then let the findings fill the room under it.