From Threshold to Test
Core challenge
An experiment at the Hawthorne Works in the 1920s raised a factory's lighting. Productivity
rose. Researchers lowered the lighting. Productivity rose again. Some of the original findings have since been
challenged, but the core problem hasn't dated: you can't isolate a change from the system it enters. The harder
question isn't whether a change worked. It's how you'd know, while you're inside it, whether you're seeing the change
itself or just the fact that something changed.
The case study
A recent session put us inside that problem using Coca-Cola's 1985 decision, releasing the same
information the real board received, stage by stage.
Stage 1. Pepsi's blind taste tests, and our own, showed a clear preference for Pepsi. Market share was slipping, and customer trust was eroding. The diagnosis looked obvious: build a new recipe, fast.
Stage 2. Internal tests on two hundred thousand people showed overwhelming preference for the new recipe. Here we split from history. Instead of full replacement, we chose a side launch in a small region, a two-way door we could reverse. We didn't define any rule for when to actually walk back through it.
Stage 3. Complaints came in fast. Daily calls rose from 400 to 1,500, and on peak days to 8,000. More than 40,000 letters arrived. People stockpiled the original formula and organised boycotts. Most of us argued to continue: the complaints were a small share of the market, adjustment takes time, people don't always know what they want until they've lived with a change. Only one person argued for immediate reversal.
We built a two‑way door and then forgot to decide when to use it.
We had a door and no rule for using it. That gap is what made the exercise resemble the
original case. Coca-Cola's own reversal is usually traced to a decision to wait and see how the first weekend of
July's sales performed, a checkpoint assembled after the crisis was already underway, not fixed before launch.
We had applied the two-way-door principle and still drove in the same direction as the
original executives. Coca-Cola later said it had underweighted brand attachment and the psychological cost of
replacement. We saw that exact risk directly, during the session, and still failed to reverse. Both the original
board and our hypothetical board were using the same complaint numbers. Nobody had agreed beforehand what those
numbers would need to show to count as failure.
A second case: the reference point wasn't the problem
We then turned to more common, relatable examples, starting with a newly assembled development
team. Someone proposed judging performance by hours logged against the original estimate: falling short of the
estimate indicating that the team was underperforming.
The first objection was that the estimate itself was a guess made before anyone knew the team's
pace and tech debt. That's fixable with more data. The second objection went further: even a well-calibrated estimate
can't tell you whether early friction is the normal mess of a team finding its rhythm, or a lasting problem.
A missing comparison point and a system not yet stable enough to compare are two different problems.
The issue wasn't the reference point anymore. It was whether the team had been together long
enough for any comparison to mean something. A team mid-formation isn't a stable object you check against history.
It's a trajectory, read by watching it move, not by freezing one point and judging whether that point looks acceptable.
A third case: the question underneath both
Two people were asked to imagine moving to a new country. Same facts, opposite conditions for
calling it a failure. One wanted novelty, and would read a quiet, predictable life as proof the move hadn't worked.
The other wanted a specific opportunity, and would read that same quiet life as fine, the failure being that the
opportunity never showed up.
Unlike the previous samples, more data wouldn’t settle this. Nothing was missing from the
evidence. What was missing was a criterion, decided before the outcome, for what the move was actually supposed to
achieve.
The order we got it in
We started the evening trying to interpret results, first hunting for a better number, then
finding out a good comparison means nothing if we don't understand the dynamics.
Only at the end did we see both questions depend on something earlier: knowing what the intervention is for, clearly
enough to say what result counts as success, what deviation matters, and when we've waited long enough for a result to
mean anything.
The order we figured this out in ran almost backwards from how the pieces depend on
each other. We reached for a threshold before we had a baseline. We questioned the baseline before asking
whether enough time had passed to measure anything. Only then did we hit the question that gives the measurement its
meaning in the first place. That's why these decisions get hard after the fact. Once a result is in front of you, the
temptation is to argue about what the number means.
The harder work is deciding, before the number exists, what you're going to let it mean.