In this guide
How Many Visits Is Enough
The number of visits you need is driven by how many outlets you run and how confidently you want to separate a weak site from a weak morning, and for most networks the honest floor is several visits per outlet per cycle rather than one. A single visit to a single outlet produces an anecdote. It captures one member of staff, at one hour, on one day, handling one scenario, and any of those four can be unrepresentative without anybody lying. Repeat the visit and the noise begins to average out; repeat it across outlets and you can start comparing sites to each other rather than to an impression. Small networks can be covered completely, and should be, because sampling a network of twenty outlets saves very little and forfeits the comparison. Large networks have to sample, and the sample is then built to cover formats, regions and trading patterns rather than to be evenly spread.
Sampling by Outlet Count
How many outlets you run changes the shape of the programme more than any other factor. Small networks are best covered by something close to a census. Where there are twenty or thirty outlets, sampling saves very little and forfeits the thing the programme is most useful for, which is comparing sites against each other. Visiting every outlet, several times, is affordable at that scale and produces a dataset where every site can be discussed on its own evidence. Large networks have to sample, and the sample is then constructed rather than drawn evenly. It has to represent the formats you operate, the regions you trade in, the age profile of the estate and the range of trading volumes, because a sample that over-represents metropolitan flagships will describe those and nothing else. Population size matters less than people expect, and this trips up buyers repeatedly. The number of observations needed to support a conclusion at a given confidence level rises very slowly once the population is reasonably large, so a network of six hundred outlets does not need twice the visits of one with three hundred to say something at the network level. What scales is the per-outlet conclusion, not the network one.
Frequency Against Depth
A fixed budget forces a choice between visiting fewer outlets more often and visiting more outlets once, and the two answer genuinely different questions. Fewer outlets more often detects change. With repeat visits at the same sites you can see whether performance is moving, whether a corrective action worked, and whether a bad result was an incident or a pattern, because you have a series rather than a point. What it cannot tell you is how the rest of the estate is doing. More outlets once measures level. It establishes where the network stands as a whole, identifies which sites sit at the extremes, and gives you a baseline against which everything later is compared. What it cannot do is distinguish a poor site from a poor morning, because a single observation carries both and there is no way to separate them. Most programmes need the second in year one and the first thereafter. Establish the level across the whole estate, find the outliers at both ends, then concentrate repeat visits where the answer actually matters, releasing budget from the sites that have already demonstrated they are stable. Reviewing that allocation at the end of every cycle is what keeps the programme answering a current question rather than repeating the one it was designed around.
Weighting by Risk and Value
Once the baseline exists, spreading visits evenly is the least useful allocation available. Three weightings earn their place. New outlets and new managers come first, because performance in the opening months is both most variable and most correctable, and a site that establishes bad habits early keeps them. A visit schedule that treats a store open six weeks the same as one open six years is spending in the wrong place. Previous poor scores are the second weighting, and they need care. Revisiting a weak site more often is right, but the purpose is to see whether it has improved, which means the repeat visit has to be far enough after the intervention to be fair and close enough to be actionable. Revisiting immediately measures the intervention, not the site. Revenue concentration is the third. Where a small number of outlets carry a large share of turnover, a compliance or service failure at one of them costs more than the same failure at a marginal site, and the visit allocation should reflect that exposure rather than treating every outlet as an equal unit of the estate.
What the Numbers Can and Cannot Support
A visit programme supports conclusions at two levels, and the two need very different amounts of data. Network-level conclusions are the easier case: with visits spread across enough outlets, the aggregate score genuinely describes how the network is performing, and movement in that aggregate between cycles is meaningful even though no individual outlet was visited often. Outlet-level conclusions are much more demanding, because a score for one site rests on however many visits that site received. To say that one outlet is performing worse than another, rather than that it had a worse morning, needs repeat visits at that outlet across different days, times and staff. Where an outlet has been visited once, the honest statement is that on one occasion these things were observed, and nothing more. A score stops being evidence of anything at the point it is used for a decision the sample cannot carry. A single-visit score driving an individual's appraisal is the clearest example, since it attributes to a person a result that may belong to the hour, the queue or the day.
Setting the Cycle
A defensible first-year plan does two things at once: it establishes a baseline for the network and it gives each outlet enough visits that its own score means something. Those pull in opposite directions on a fixed budget, and the usual resolution is to cover every outlet at a modest depth in the first cycle rather than a subset deeply, because the first year's job is to find out where you actually stand. Spread the visits across days, times and trading conditions rather than clustering them, since a cycle run entirely on weekday mornings measures weekday mornings. After the baseline, redistribute. Outlets performing consistently well need fewer visits to stay confident about them; outlets that scored badly or moved sharply need more, and the released budget funds that. Reviewing the allocation each cycle is what keeps the programme measuring something useful instead of repeating itself. Bringing in a provider to design the cycle is worth it where the network is large, where formats differ enough that one scorecard will not serve, or where the results will drive incentives, and Mystery Audit / Mystery Shopping covers how a first cycle is scoped.
