
Here’s what it fixes
A director asks for a configurable dashboard. It’s not the first time, and she’s convincing: her team of claims handlers needs “more visibility.” It’s a clear, popular request, so it lands right at the top of the backlog.
In a team running real discovery, it lands somewhere else first: the hypothesis log. “We believe claims handlers want a configurable dashboard, because they’ve asked for more visibility. If we ship one, we expect daily active use above 70%. We’ll know we’re wrong if handlers can’t articulate what they’d put on it.”
A few days and five interviews later, it’s dead. Not one handler named a widget they’d actually put on the board. All five want something else: the age of their work queue. The dashboard never gets built. The number ships in the next increment as a new field on a page, and the handlers get what they really needed. The team saves weeks of wasted effort delivering something nobody would have used.
That log entry earned its keep because it was written to lose. Most entries aren’t, and that’s why most hypothesis logs don’t work.
If you’re new, welcome to Customer Obsessed Engineering! I write about an article a week. Free subscribers can read roughly half of most articles, plus my free articles. Learn more.
What makes an entry a hypothesis
I’ve made the case for continuous discovery (weekly customer contact, the economics of killing bad ideas while they’re cheap) in Continuous discovery: the missing half of delivery. This piece is about the log itself, because the log is where discovery becomes real or fake.
First, a hypothesis is not a feature request, and putting “we believe” in front of one doesn’t change that. Just, “we believe users want batch upload” is a wish — and it’s still a feature request, because it can’t lose. Any interview can be read as support; anyone’s lukewarm reaction can be explained away. That's how, six weeks later, a feature ships because it was always going to ship.
A hypothesis is a claim about the world that could turn out to be false, written so the team knows in advance what evidence would change its mind. Ask one question of every entry in your log: what will kill it? If there’s no answer, you aren’t testing a hypothesis, and you aren’t running discovery: you’re just checking boxes on a delivery team.
Writing the hypothesis
The playbook’s hypothesis log template uses a five-part format:
We believe [who] experiences [problem] because [assumed cause]. If we [solution], we expect [measurable outcome]. We’ll know we’re wrong if [disconfirming evidence].
Every clause pulls its weight. Naming who rules out “users.” If you can’t say which persona or segment has the problem, you’ve found your first gap before running the test. Separating the problem from the assumed cause matters because the cause is usually where you’re wrong: the handlers’ visibility problem was real; the assumption that a dashboard would solve it was not. And the expected outcome needs a number so you can’t renegotiate it after the fact into whatever the test happened to produce.
Then the last clause: the disconfirming evidence. This is the most important, and also the one teams skip all the time. Stating up front what would prove you wrong makes the entry falsifiable, the same line Karl Popper drew between science and everything that dresses like science. It turns the entry from a deliverable into a test, and if writing it feels uncomfortable, great! The discomfort is the point. An entry without it is a milestone… not something that feeds evidence and informs what to do next.1
One more rule: one claim per entry. Don’t conflate things. Pulling lots of levers at once doesn’t tell you anything useful. Compound hypotheses can’t be settled. When half of one validates, the whole entry limps forward on partial evidence and nobody can say what was learned.
This newsletter grows by word of mouth… I’d really, truly appreciate it if you could refer a friend. Your referrals make it worthwhile.
Source, risk, test and cost
The rest of the entry is four short fields, governed by one mechanic: the experiment runs before anything gets built. The log isn’t a list you tick off as delivery proceeds. It’s a log of experiments to run, and every claim gets proven or disproven before the work behind it goes any further. Treat it as a gate. The playbook enforces exactly that in the definition of ready for increment selection: it requires current discovery evidence before you choose what goes into an increment.

