amplifyWeb
← Back to the blog
·4 min read

Most A/B Tests Are Opinions With a Confidence Interval Attached

Here is the uncomfortable part of conversion work that nobody puts on a slide: the majority of tests that get built were chosen because someone in the room felt strongly, not because the evidence pointed there. The statistics come later. They get bolted onto a decision that was already made in a meeting. The confidence interval makes the opinion look rigorous. It does not make the opinion right.

I want to make the case for a different starting point. Not "what should we test?" but "what has earned the right to be tested?"

The problem with idea-first testing

Idea-first testing feels productive. You brainstorm, you fill a backlog, you ship. The trouble is that a backlog built on enthusiasm has no natural filter. The loudest voice, the most recent conference talk, the thing a competitor just did: all of it lands in the queue with equal weight. Then you spend real traffic finding out which hunches were wrong.

Traffic is the most expensive thing you have. Every visitor routed into a weak variant is a visitor you cannot get back. Idea-first testing treats that traffic as a brainstorming tool. It should be the last resort, not the first.

Evidence-first testing

The alternative is to refuse to write a single variant until independent evidence lines up. I use three streams, and I do not proceed unless they agree.

The first stream is behaviour. Where does the funnel actually leak? Not where you assume it leaks. A step-by-step view of where people abandon tells you where the money is, and it is almost never where the debate in the room was focused.

The second stream is the market. What are competitors spending real money to say, over and over, month after month? Sustained ad spend is a confession. It tells you which messages are working well enough to keep paying for. If the market keeps hammering a benefit that your page barely mentions, that is a signal you did not have to guess at.

The third stream is history. Has this idea, or something close to it, been tried before? What happened? A test you ran last year is worth more than a case study you read about someone else's audience. If you are not keeping a record of what you have already learned, you are paying full price for the same lesson twice.

What "agreement" looks like

An idea that survives all three streams looks like this: the funnel shows a specific drop-off, the market is loudly addressing the exact concern that drop-off implies, and your own history shows that addressing that kind of concern has moved the needle before. That is not a hunch. That is a conclusion three independent sources arrived at without talking to each other.

When the streams disagree, the idea is not dead. It is unfinished. Maybe the market signal is strong but the funnel does not back it up, which means the problem is real but lives somewhere else on the page. Reworking an idea on a spreadsheet is free. Reworking it in live traffic costs you a test cycle and a chunk of your credibility.

Why this matters beyond any single test

A program built on evidence-first selection does something a hunch-driven program never does: it compounds. Every test result feeds back into the history stream, so next quarter's decisions are sharper than this quarter's. A program built on opinions stays exactly as smart as it was on day one, forever, because it never files anything away.

The confidence interval was never the problem. The problem is treating it as proof of a decision that evidence never actually made. Put the evidence first, and the statistics stop being a costume. They become what they were supposed to be: the final check, not the opening argument.