Changes you cannot split-test
Experimentation is the gold standard and does not apply to everything. You cannot give half your customers a different depot network, run two rota structures in one team, or trial a pricing model on half a market without contaminating the other half.
The alternatives are to guess, to rely on a supplier's claims, or to build a model of the system and try the change in it. The third is usually the best available evidence.
What a useful simulation contains
- The entities that matter - orders, vehicles, staff, machines - and their real distributions from your data.
- The rules the system follows, including the informal ones people actually use.
- The constraints that bind - hours, capacity, opening times, break rules.
- The randomness, drawn from observed variation rather than assumed.
- A validation step: reproduce a period you already know the outcome of.
That last step separates a useful simulation from an expensive opinion. If it cannot reproduce last quarter reasonably, its view of next quarter is not evidence.
Compare options, do not predict absolutes
A simulation's absolute numbers carry all the error in its assumptions. Its comparisons are far more reliable, because the assumptions affect both options similarly.
So the right question is 'is option A better than option B, and by roughly how much' rather than 'what will our costs be'. Presenting simulation output as a forecast invites a precision it does not have and damages credibility when reality differs.
| Question | Simulation suitability |
|---|---|
| Which of three depot layouts is best? | Good |
| Roughly how much would layout B save? | Reasonable, as a range |
| Exactly what will costs be next year? | Poor |
| What breaks if volume rises 30%? | Very good |
Stress testing is where it shines
The most valuable output is often not the recommended option but the failure modes. Pushing volumes up, removing a vehicle, extending a lead time - these reveal where a plan breaks and how suddenly.
Systems frequently degrade non-linearly: fine at 90% utilisation, unworkable at 95%. Finding that cliff before you are on it is worth more than a marginal efficiency gain, and it is very hard to discover any other way.
Be honest about the assumptions
Every simulation embeds assumptions, and the ones that matter most are usually about behaviour - how customers respond to a longer lead time, how staff adapt to a new rota. Those are the hardest to justify and the most influential.
List them explicitly, test how sensitive the conclusion is to each, and report the ones that change the answer. A conclusion that holds across a wide range of a doubtful assumption is much stronger than one that depends on getting it right.
If the simulation cannot reproduce last quarter, it is not evidence about next quarter.