Most non-profit leaders meet evaluation as a demand, not a tool. A funder asks for it, a board wants it, and someone scrambles to produce a report that proves the program worked. That framing does real damage, because evaluation is not one thing you do at the end to prove yourself; it is a family of related methods, each built to answer a different question at a different point in a program's life. Reach for the wrong one and the cost is not only a wasted budget line. You can distort a program that was still finding its shape by judging it too soon, or learn nothing useful from one that has already settled by treating it as though it were still an experiment. The choice of method is really a choice about timing, and timing is the part most organizations skip.
Two ways to time an evaluation: steer as you go, or judge at the end
01 · Three questionsThe question decides the method
It helps to stop thinking of summative, formative, and developmental evaluation as a maturity ladder, with developmental as the casual, less rigorous option and summative as the serious one. They serve different decisions, and none is a lesser version of the others. Summative evaluation asks: what judgment can we make? Formative evaluation asks: what should we improve? Developmental evaluation asks: what is taking shape, and what are we learning as it changes?
The evidence: current government evaluation guidance starts with the decision and the questions, then selects an approach that fits the intervention, its context, and how the findings will be used.12 That matters because the three labels are not sealed boxes. A mature program can need formative feedback on delivery and a summative judgment on outcomes at the same time. A new program can still owe funders a clear account of what it did, even when an impact estimate would be premature.
A rough way to tell where you are: has the program been delivered the same way, by the same kind of team, more than once, with participation that has settled into a pattern? Or is the model still changing month to month because you are still learning what it even is? The honest answer to that question, asked before you pick a method, saves most of the trouble that follows.
For practitioners: telling which stage a program is actually in +
Three rough signals help without much guesswork. Developmental territory: the model still changes noticeably month to month, key assumptions have not been tested yet, and the people running it still disagree about how it is supposed to work. Formative-ready: the program has been delivered consistently for at least one full cycle, roles and referral pathways have settled, and the core logic is not seriously contested, even if the details need tuning. Summative-ready: the program has run the same way across multiple cycles or sites, some form of baseline or comparison exists, and the people involved broadly agree on what success would look like.
When in doubt, assume you are earlier than you think. Programs are declared "stable" for funding reasons more often than they are declared stable because they actually are, and a premature summative report tends to punish the program for the mismatch rather than reveal anything true about it.
02 · SummativeJudging a finished thing
Summative evaluation supports a judgment against stated criteria or outcomes. It often happens near a decision point, but it does not have to wait until a program is over. Organizations, communities, and funders all need an honest accounting of whether something worked, for whom, under what conditions, and at what cost. The trouble is not judgment; it is treating judgment as the only purpose, or reaching for a causal verdict before the model, outcome, and comparison are credible. Used well, summative work informs a real decision: renew, scale, redesign, or stop. Used alone, it can become an autopsy: precise, honest, and delivered after the point where anyone could have acted.
Good summative work still requires more than a before-and-after number. It needs a credible sense of what would have happened anyway, some comparison or baseline, or the result is just a description of the world, not evidence that the program caused it.
03 · FormativeTuning what is basically sound
Formative evaluation feeds findings back while there is still time to adjust. It usually assumes enough of the model is stable to improve delivery, but it can also expose a theory that needs reconsideration. For most established programs, this is the workhorse, and it is badly underused. Organizations tend to skip straight from launch to proof, and never build the feedback loop that should sit in between.
In practice, formative work looks unglamorous: short cycles, a few weeks rather than a year; frontline staff describing what is actually happening, not just what the plan said should happen; small adjustments to sequencing, dosage, or staffing rather than a redesign. The discipline is keeping it separate from judgment. The moment a formative check starts producing a pass or fail score, staff stop being honest in it, and you lose the only tool built for catching problems while they are still cheap to fix.
For practitioners: running a formative loop without it turning into a summative one by accident +
Keep the cycle short, weeks rather than a year, so the feedback still arrives while the team can act on it. Report first to the people who can actually change something, frontline staff and the program lead, before it goes up to a board or funder; a formative finding that reaches decision-makers only after it has been filtered for a funder audience has usually lost its usefulness. Use plain descriptive language, on track, needs adjustment, worth watching, rather than a pass or fail score, since scoring is what turns a formative conversation defensive.
Separate the learning conversation from the funding conversation wherever you can. The moment staff suspect a formative check will affect next year's grant, they will describe the program they wish they were running instead of the one that actually exists, and the whole exercise stops working.
04 · DevelopmentalWalking alongside something still being invented
Developmental evaluation, associated most closely with Michael Quinn Patton, is for work that is genuinely innovative or complex, where the model is not settled because it cannot be yet: an untested approach, a context that keeps shifting, several partners who do not fully agree on what the problem even is. It does not wait for a finished thing to judge. It walks alongside the work, feeding back patterns as they emerge so the team can adapt the design itself, not just the delivery.3
An honest limit: "complex" cannot become a synonym for "we do not want to be judged." Developmental work still needs explicit questions, documented decisions, evidence strong enough for the claim being made, and a point at which the evaluation architecture is reconsidered. Complexity changes what rigour looks like; it does not remove the obligation.
This is the real difference from formative work. Formative evaluation tunes a model that everyone agrees on. Developmental evaluation helps a team figure out what the model even is, and it tracks the reasoning behind each shift, not only what changed, so that when the work does stabilize, the logic behind it is not lost.
Before you choose a method, choose your question: are you inventing something, improving something, or accounting for something?
Here is that mapping laid out plainly.
Which evaluation, and when: match the method to where the program actually stands
05 · The mismatchProof demanded too early
The mistake I see most often is a mismatch between the question and the tool: a funder attaches a summative requirement to the first year of something genuinely new, and the team spends its scarce energy trying to prove outcomes that could not possibly have stabilized yet. The program looks like it is failing, when what actually failed was the timing of the question. Worse, a team under that pressure will sometimes quietly narrow the program until it produces a number that is easy to defend, which is its own kind of harm.
The honest move is to negotiate, and to do it before the method is fixed rather than after. Propose a developmental or formative approach for the period while the model is still forming, and agree on a concrete point, a cohort, a year, a specific stability milestone, at which summative measures become appropriate. Bring that plan to the funder rather than waiting to be told what to report. Funders are more open to this than people expect, because most of them would rather fund real learning than a performance of certainty.
For practitioners: negotiating evaluation timing with a funder +
Bring a two-phase plan to the funding conversation rather than waiting to be handed a method: a developmental or formative approach while the program is still forming, summative once it has genuinely settled, with the switch point named as a specific milestone rather than left vague. Offer a lighter interim report in the meantime, a short account of what has been learned and adjusted, so the funder is not left with nothing while the program matures.
Frame the request as protecting their investment, not avoiding accountability. A funder who forces a premature verdict on a young program is more likely to defund something that would have worked than to catch something that would not have. Naming that risk plainly, in their terms, tends to land better than naming it in yours.
06 · Whose definitionSuccess is not a neutral question
None of these methods is culturally neutral, and the summative instinct in particular can quietly impose a definition of success that the community being served never chose. A target fixed before anyone consulted the people the program serves, then checked at the end by summative evaluation, only ever tells you whether the program hit somebody else's mark.
Developmental and formative approaches, because they happen earlier and repeatedly, leave more openings for the people a program serves to shape what "working" actually means, and for that shaping to change something before the story is already written. That is not a soft or secondary consideration. It is the difference between measuring the right thing and measuring the convenient one, and it is worth building into the evaluation plan from the very first conversation, not adding in afterward.
07 · Try itChoose an evaluation architecture, not just a label
The chooser below recommends a lead approach and a companion approach. That is deliberate. Real evaluations often need one method to steer the work and another to answer the accountability question.
Evaluation approach chooser
Answer for the program as it exists now, not the version described in the proposal. This is a planning aid; it does not select a research design or determine causal validity.
None of this is complicated once you separate the three questions. What is hard is resisting the pull toward summative language before a program is ready for it, and resisting the pull toward calling a settled program still "developmental" because that framing is more comfortable. Match the method to the moment, name that moment honestly, and the rest of the evaluation gets much easier to get right.
Sources and method note
- HM Treasury. (2026). Magenta Book: Central Government Guidance on Evaluation. Evaluation questions, intended use, intervention characteristics, context, and complexity should shape design.
- Treasury Board of Canada Secretariat. Evaluation 101 Backgrounder. Logic models, evaluation matrices, questions, indicators, data, analysis, and limitations.
- Patton, M. Q. (2011). Developmental Evaluation: Applying Complexity Concepts to Enhance Innovation and Use. Guilford Press.
- HM Treasury. (2026). Handling Complexity in Policy Evaluation. Evaluation in complex systems may need regular review, adaptation, and theory-based reasoning.