It is the most common confusion in the non-profit world, and it hides in plain sight inside almost every annual report. An output is what you did: workshops held, people served, reports produced. An outcome is the change that resulted: people who are healthier, decisions that improved, a community that is safer. Funders ask for outcomes. Organizations report outputs. The gap between the two is where a lot of money quietly fails to do what it was meant to, and the report can look completely successful the whole time it is happening.
Every output should end in a change someone can feel, or it is just an activity that happened
01 · The confusionThe pull toward outputs
Outputs are easy to count and entirely within your control. You can guarantee you will run the sessions, print the pamphlets, answer the calls; the number at the end of the year is simply a tally of things you already decided to do. Outcomes are harder: they take longer to appear, they are shaped by forces outside the program, and you can never fully claim them as yours alone. So the temptation, especially under deadline, is to retreat to what you can count and let "we trained 400 staff" quietly stand in for "practice actually changed." It does not, and a report full of confident outputs can describe a program that helped no one and still read as a success.
None of this makes outputs worthless. You have to track what you did, who you reached, and whether delivery occurred as intended. Implementation research distinguishes outcomes such as adoption, feasibility, fidelity, cost, reach, and sustainability from service and participant outcomes.3 The problem is not measuring outputs; it is stopping there and calling them evidence of a change they cannot prove.
02 · The chainThe "so that" test
The discipline that fixes this is almost embarrassingly simple. For every output, ask "so that what?" and keep asking until you reach a real change in someone's life. We ran the workshops, so that staff understand the new approach, so that they use it with families, so that families experience better care. Each link is a causal claim that needs evidence; together they form a theory of change, whether or not anyone ever drew the diagram. Current evaluation guidance recommends making those links and their assumptions explicit before selecting measures.12
Try running the test backward, too. Start from the outcome you actually want and ask what has to be true just before it for that outcome to occur, and what has to be true before that. If your current output does not sit somewhere on that chain, it may be a busy, well-intentioned activity that is not actually connected to the change you are claiming.
If your report could be equally true whether or not the program helped a single person, you are measuring outputs.
For practitioners: running the "so that" test on your own logic model +
Take any line from your last report and ask "so that what?" out loud, in the room, with the people who wrote it. The first answer is usually still an output in disguise: "so that people attended," "so that materials were distributed." Keep asking. You are looking for the point where the answer describes a person doing, feeling, or deciding something differently than they would have otherwise. That is usually two or three "so that"s deeper than where the sentence started.
Do this exercise with frontline staff, not only with whoever writes the funder report. The people closest to the work usually know exactly where the chain gets shaky; they are just rarely asked, because the report gets written from the output data that is already sitting in a spreadsheet.
03 · Where it breaksThe gap in the middle
Most organizations can state the first link, what we did, and the last one, the change we ultimately want, without much trouble. The chain usually breaks in the middle, at the step nobody measures because it is the hardest one: did the people who received the activity actually change their behaviour because of it? Training completed is an output. A worker using a new approach with an actual family, next week, under real caseload pressure, is the outcome that training was supposed to buy, and it is the step almost nobody checks.
This middle step is where good programs quietly fail without anyone noticing, because the output at one end and the aspiration at the other both still look fine on paper. Checking the middle is usually cheaper than people assume: a short follow-up conversation, a supervisor's observation, a simple practice log, rather than a full outcome study.
For practitioners: measuring the middle step without a full outcome study +
You do not need a research budget to check the step everyone skips. A short structured follow-up, a handful of questions asked the same way each time, of a sample of the people who received the activity, will usually tell you whether anything changed in how they act, not just whether they liked the session. A supervisor's brief observation of practice, done occasionally and consistently, does the same job for staff-facing programs.
The point is proportional rigour: evidence strong enough for the claim you intend to make. A small, honestly described sample can reveal whether practice changed; it cannot by itself establish a population effect or rule out every alternative explanation. Match the sentence in the report to what the evidence can carry.
Drawn as a chain, the shape of that failure is easy to see.
The chain rarely snaps at the ends; it snaps in the middle
04 · The nearest outcomeMeasuring short of the final outcome
A common objection is that real outcomes, a healthier community, a safer city, take years and resources no small organization has. That is true, and it is not a reason to retreat to outputs. Pick the nearest honest outcome you can actually observe on your own timeline and budget, even a modest one, rather than the most impressive one on the chain. A shelter cannot easily prove it reduced regional homelessness this year, but it can measure whether the people it housed were still housed ninety days later, which is a real outcome, several links closer than "beds provided," and well within reach.
The goal is not to reach the most distant, most impressive outcome on the chain. It is to move at least one honest step past the output, and to say plainly which step you reached and which ones you are still assuming.
For practitioners: picking the "nearest honest outcome" when the real one is years away +
List every link on your chain from activity to the distant outcome you actually care about. Find the first link past your current output that you could plausibly measure within your existing timeline and budget, without a special grant for evaluation. That is your nearest honest outcome; report it as what it is, a real but partial step, not as proof of the distant goal.
Say explicitly, in the report itself, which further links remain assumptions rather than measured results. That sentence costs you nothing in credibility and gains you a great deal of it; funders trust an organization that names what it has not yet proven far more than one that implies it has proven everything.
05 · The costA report that could be true either way
Here is a fast test for your own reporting: read a paragraph back and ask whether it would say the same thing in a year the program quietly failed. If yes, no matter how confident it sounds, you are describing outputs, not outcomes. That test is uncomfortable to apply to your own writing, which is exactly why it is worth applying.
The cost of skipping this is not just an inaccurate report. It is a program that can drift for years, fully funded and fully busy, without anyone, including the people running it, noticing that it stopped helping the people it was built for. Outputs will not tell you that. Only an honest outcome, however modest, will.
06 · The honest counterargumentOutcome language can overclaim too
An outcome statement is not automatically evidence. "Families were stronger" may sound more meaningful than "four hundred people attended," but without a defined measure, a timeframe, and a credible account of what else may have produced the change, it is only a better-written aspiration. Organizations can move from output theatre to outcome theatre without becoming more accountable.
The other danger is erasing implementation. A program may show weak outcomes because the underlying idea failed, because it never reached the intended population, because staff could not deliver it under real conditions, or because the outcome was measured too early. Output and implementation evidence help distinguish those explanations. The honest report needs both: what was delivered and what changed, with the causal bridge between them exposed rather than implied.
07 · Try itBuild the chain your next report should test
Use the builder to move one real program statement past activity. The result is a draft measurement brief, not proof that the links are true.
Nearest honest outcome builder
Write plain language. Specific beats impressive. Your entries remain in this browser.
None of this requires a research department. It requires the discipline to ask "so that what" one more time than feels comfortable, and the honesty to report the answer even when it is a smaller, harder-won outcome than the one on your letterhead. That is the difference between a report that describes activity and one that earns trust.
Sources and method note
- HM Treasury. (2026). Magenta Book: Central Government Guidance on Evaluation. Theory of change, causal chains, assumptions, context, and evidence strength.
- Treasury Board of Canada Secretariat. Evaluation 101 Backgrounder. Inputs, activities, outputs, outcomes, indicators, methods, analysis, and limitations.
- Proctor, E., Silmere, H., Raghavan, R., et al. (2011). Outcomes for implementation research. Distinguishes implementation outcomes from service-system and participant outcomes.
- Treasury Board of Canada Secretariat. Evaluation Guidebook for Small Agencies. Logic models and the question, "What outcome resulted from the output?"