A pilot goes well, a funder gets excited, and the pressure to scale arrives almost immediately: take it to three more sites, put it in next year's strategic plan, build the case for growth funding. It is the right instinct and often the wrong sequence, because it skips past the one question that actually decides whether growth helps or hurts. Is the thing in front of us ready to grow? Scaling an unready program does not spread its success; it multiplies whatever was still fragile in it, usually with more money attached and more people watching it happen.
Readiness builds gradually. The decision to scale often does not.
01 · The instinctGrowth pressure, and the question it outruns
The rush to scale is not a character flaw. It is a structural pressure. Funding cycles reward growth narratives more than steady-state ones, a board wants to see momentum it can point to, and a program team that has worked hard for two years understandably wants to see that work reach more people. None of that is unreasonable. The problem is that "it worked" and "we know why it worked" are two different claims, and only the second one tells you much about what will happen at three times the size.
The evidence: scale-up research treats expansion as a deliberate process, not the automatic next chapter of a successful pilot. A structured scalability assessment asks about the evidence of effectiveness, costs and benefits, fidelity and adaptation, reach and acceptability, workforce, infrastructure, and sustainability before recommending expansion.2 The World Health Organization's current public-health guidance describes scale as a continuing cycle of exploring, adapting, and learning, not a one-time replication decision.1
My argument: a pilot tells you the idea can work somewhere, under some conditions, with some people. It does not automatically tell you which conditions were load-bearing and which were incidental. Readiness work is the discipline of finding out before you commit the budget and the reputation to an answer you have not checked.
02 · Three lensesThe intervention, the organization, the system
The evidence: the three lenses below condense a longer implementation-science tradition. WHO and ExpandNet distinguish the innovation, the organization expected to adopt it, and the wider environment, alongside the resource team and the scale-up strategy.3 I use three here because they are memorable enough to survive an actual decision meeting.
My framework: read readiness through the intervention, the organization, and the system. A program needs all three; strength in one does not erase a gap in another.
The intervention. Is the model actually defined and stable, or is it still a talented team improvising well and calling the improvisation a program? Is there real evidence it works, or a good first year and a lot of goodwill? You cannot replicate what you cannot describe. If the model lives mostly in one coordinator's head and habits, scaling will dilute it past recognition before anyone notices it happening.
The organization. Does the organization that will deliver the program at scale have the capacity, the leadership attention, and the back-office systems to carry it, not just the mission alignment to want to? Programs fail in expansion far more often because the host could not absorb the growth than because the idea was wrong. Finance, HR, data systems, supervision ratios: unglamorous, and usually where it actually breaks.
The system. Will the surrounding environment, funders, partners, referral sources, policy, community trust, support the program at a larger footprint, or was the pilot's success partly a product of conditions that will not travel? A program that thrived because one relationship protected it from the usual friction is not yet ready for a version of itself without that relationship.
Readiness is not a brake on ambition. It is what lets ambition survive contact with reality.
For practitioners: turning three lenses into an actual instrument +
The three lenses are easy to agree with and easy to apply loosely. To make them useful, turn each into a short, specific set of questions you actually ask, and ask them of more than one source. Interview the program team, yes, but also frontline staff who were not in the room when the model was designed, and at least one external partner who depends on the program working. The gap between how leadership describes the model and how a frontline worker actually executes it on a Tuesday is usually the most honest readiness signal you will get.
My recommendation: if local circumstances force you to weight one lens more heavily, weight the organization. A strong model inside a fragile host still fails; a decent model inside a well-run, well-resourced organization has a chance to improve as it scales. That is practitioner judgment, not a universal empirical rule. Make the weighting visible so others can challenge it.
03 · Built, not assumedReadiness, stage by stage
It helps to borrow a lens from individual change work, because organizational readiness follows a similar logic. A person does not adopt a change just because it was announced. They move through being aware it is coming, wanting it, knowing how to do it, being able to do it, and then having something reinforce the new way of working once the novelty wears off. Skip a stage and the change stalls, no matter how good the plan looked on paper.
An honest limit: ADKAR is the practitioner framework I use; I do not treat it as a validated scale-readiness instrument. The research literature defines organizational readiness more narrowly as a shared resolve to implement a change and a shared belief that the organization can carry it out.4 ADKAR adds a useful sequence for thinking about the human work, but the sequence should organize inquiry, not substitute for evidence.
Scale is the same kind of change, aimed at an organization instead of a person. Awareness that growth is coming is usually the easy part, and so, often, is desire, because growth is exciting and people like being part of something that is working. What gets skipped is the back half. Does everyone who has to deliver the program at a larger size actually know how, in enough operational detail to do it consistently? Can they do it given their real caseloads, training, and tools? Is there something, supervision, data review, a community of practice, holding the new pattern in place once the launch excitement fades? Most scale decisions check the first two stages and assume the last three will sort themselves out. That assumption is usually where the fall happens.
Picture that sequence as five bars of different heights: the tall ones do not matter if the shortest one is not yet built.
The change stalls at whichever stage is weakest, not at the average
04 · The false positiveA good pilot that fails at scale
Pilots can succeed for reasons that do not scale. They may run with an unusually committed staff member or two, a local champion clearing obstacles that will not exist at site twelve, participants who opted in with extra motivation, and a level of attention that comes automatically from something being new and watched. None of that is a flaw in the pilot, and none of it is deception; it is part of what a pilot is designed to do, create enough protection to learn whether an idea is promising.
The risk is mistaking the protection for the model. A readiness assessment tries to separate the two: what part of this success travels with the design, and what part was actually the conditions around it? The honest version of that question is uncomfortable, because it can mean telling a team that worked hard and got real results that the result may not be theirs to keep once the champion moves on or the extra attention goes away. Asking it anyway, before the expansion, is far kinder than letting site twelve find out.
For practitioners: separating the model from its conditions +
A useful test: ask what would have to be true for this program to work without its founding champion, in a location the champion has never visited, run by staff who were not there for the original enthusiasm. If the honest answer is "not much would change," the model is probably sound. If the honest answer involves a lot of silent dependence on one person's relationships or judgment calls, you have found the thing that will not travel, and it is better to know that now than after the expansion is funded.
If the pilot ran in more than one location or cohort, look hard at the variation between them rather than only the average. The sites or cohorts that struggled usually reveal what was actually load-bearing more clearly than the ones that succeeded; success can hide a lot of dependencies that only failure exposes.
05 · The uncomfortable auditThe purpose of a readiness assessment
A readiness assessment is deliberately uncomfortable. Its job is to surface the gaps while they are still cheap to fix, before they are built into a larger, more expensive, more public version of the program. Done well, it does not produce a single yes-or-no verdict. It produces something more useful: here is what has to be true before this can scale, here is the evidence for our present judgment, here is what remains uncertain, and here is a realistic order for closing the gaps.
The evidence: the Intervention Scalability Assessment Tool does not rely on one enthusiasm score. It examines ten domains across evidence, context, implementation, and sustainability, then makes the strengths and weaknesses visible before a recommendation is made.2 Implementation research also separates outcomes such as acceptability, feasibility, fidelity, cost, reach, and sustainability from the outcomes experienced by the people a service exists to help.5 Both matter. A program can improve lives in a pilot and still be impossible to deliver well at scale; it can also spread smoothly while losing the effect that justified spreading it.
That last part matters as much as the list itself, which is the subject of the next section. A readiness assessment that hands back a dozen unranked conditions is not much more useful than no assessment at all; the organization still has to guess where to start. The assessment earns its keep by turning a vague worry, are we ready, into a short, ordered, ownable set of things to do next.
06 · SequencingOrder matters as much as the list
Not every precondition can be worked on at once, and not every precondition should be. Some can run in parallel: raising expansion funding and drafting the training curriculum do not depend on each other. Others have a strict order, and getting the order wrong is its own kind of readiness failure. You cannot train new staff on a model that has not been documented yet. You cannot expect frontline ability before people have had real practice at the new version of the work, not just a slide deck about it. You cannot expect reinforcement, the systems that keep a practice alive after the launch buzz fades, to hold a pattern that ability has not actually established yet.
This is the same logic as the readiness stages themselves: each one depends on the one before it, and trying to shortcut the sequence usually just moves the failure later and makes it more expensive. A sequenced readiness roadmap, with a named owner for each item and a real trigger for reassessing rather than a fixed date on a calendar, is what turns "we think we're ready" into something an organization can actually stand behind.
For practitioners: building the roadmap without stalling everything +
Sequencing preconditions does not mean doing one thing at a time. Group them into what can run in parallel and what has a genuine dependency, and only enforce order where the dependency is real; treat everything else as concurrent, so the readiness process does not become an excuse for a year of nothing happening.
Give each precondition a single named owner, not a committee, and a plain description of what "done" looks like for that item specifically. Build in a real trigger for re-checking readiness, a pilot cohort completing, a new site's first ninety days, rather than a fixed date chosen because it was convenient for the board calendar. A readiness process that only ever gets checked once, at the start, is not really a process; it is a permission slip.
07 · The honest counterargumentReadiness can become avoidance wearing professional language
There is a real danger on the other side of this argument. An organization can keep asking for more readiness evidence until a promising program dies in committee. A large incumbent service can be allowed to continue on thin evidence while a smaller innovation is asked to prove every assumption before reaching anyone new. Communities can be made to wait for methodological comfort they did not ask for. Readiness language can become a gate, and the people holding the gate are not always the people carrying the cost of delay.
That is why the answer is rarely "ready" or "not ready." The better decision is often a bounded scale step: one additional site, one new population, or one operating condition that meaningfully tests the weakest assumption. Name what must stay faithful, what may adapt, what evidence will be collected, who gets to judge acceptability, and what result triggers the next decision. WHO's latest guidance makes adaptation and learning part of scale itself.1 Readiness should determine the size and conditions of the next step; it should not become a demand for certainty no real-world program can meet.
08 · Try itA five-minute readiness check
The questions below turn the three lenses into a first conversation. They do not produce a certification. Answer them with a mixed group if you can, then investigate the disagreements; a split between leadership and frontline answers is often more useful than the average.
Next-step readiness check
Choose the most honest answer for each condition. Nothing leaves your browser. The report treats the weakest lens as the constraint because an average can hide the one gap most likely to break the expansion.
None of this is an argument against growth. Good programs should reach more people; that is usually the point of running them at all. It is an argument for sequencing growth deliberately instead of assuming it, for treating readiness as something you build on purpose rather than something you discover you were missing after the ribbon-cutting. A pilot can show that an idea is promising under known conditions. Readiness work identifies which conditions must travel, which may adapt, and what the next stage still has to learn. That is what gives the larger version an honest chance of helping more people without becoming a thinner version of the thing that worked.
Sources and method note
- World Health Organization. (2026). Scaling innovations in public health systems: guidance and toolkit. The framework centres exploring, adapting, and learning during scale.
- Milat, A. J., Lee, K., Conte, K., et al. (2020). Intervention Scalability Assessment Tool: A decision support tool for health policy makers and implementers. Health Research Policy and Systems, 18, 1.
- World Health Organization and ExpandNet. (2010). Nine steps for developing a scaling-up strategy.
- Weiner, B. J. (2009). A theory of organizational readiness for change. Implementation Science, 4, 67.
- Proctor, E., Silmere, H., Raghavan, R., et al. (2011). Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Administration and Policy in Mental Health, 38, 65-76.