Back to blog
•6 min read

How to track whether a conversion fix worked

Comparing one audit against another is a one-group pretest-posttest, the weakest design there is. Here is the arithmetic for when its answer becomes readable, and the four things to set up before you ship the fix.

how to track conversion fix impactdid my cro change workmeasure impact of website changesconversion fix measurement

"I made the changes, the number went up a bit, and I still cannot tell you whether it was me." That is where most founders land, and it is not carelessness. Comparing one before against one after is a weak design with a name. Tracking the impact of a fix means dating every change and reading the step it was meant to move, not your sitewide conversion rate.

How do you track the impact of a conversion fix?

Write down what you changed and the date you shipped it, pick the single number that change was supposed to move, then compare matched periods either side of that date. Read the step, not the whole site. The step carries more signal per visitor.

Most people instead re-run the audit a month later to see whether the report looks friendlier. Two audits are two photographs, and photographing something twice is not a trend.

What measurement design are you actually running?

A one-group pretest-posttest. You measure, you change something, you measure again, with no control group anywhere. Research methods textbooks file that under pre-experimental rather than experimental, and they name five separate reasons its result can fool you into crediting yourself.

The five are history, maturation, testing, instrumentation and regression to the mean. Translated: something else happened, your visitors changed, measuring changed things, your tracking changed, or the number was drifting back anyway. History is the one that mauls websites. A competitor cut prices, your best affiliate went quiet, Google reshuffled your rankings. None of that is in your change log. All of it is in your conversion rate. Naming the design at least hands you a list to rule out.

Want the fix list first? Revslip ranks what to change on your URL.


Did my CRO change work, or did the week change?

Usually you cannot tell from one comparison, and the bigger the jump looks, the more suspicious you should be. Large before-and-after moves on small traffic are mostly noise finding a flattering shape, and the size of that distortion has been measured.

Kohavi, Deng and Vermeer put it bluntly in A/B Testing Intuition Busters at KDD 2022: "any figure that looks interesting or different is usually wrong." They cite work on underpowered studies where significant effects were inflated by 25% to 50%, and note that below 50% power the exaggeration stops meaning anything.

That sits awkwardly against the case-study genre, where 30% lifts are routine. Both can be true. The same authors report in their KDD 2014 rules of thumb that Bing moves by "0.1%-1% after a lot of work", because Bing has no gross defects left. A site nobody has audited usually does. Your first fix is the one big enough to see. Your fifth is not.

How big does a fix have to be before you can see it?

Big enough to clear the noise in the number you are reading, and that threshold shifts by a factor of 73 depending on which number you pick. A sitewide conversion rate needs tens of thousands of visitors per period. A single step needs hundreds.

The sample size rule of thumb from Kohavi, Henne and Sommerfield's 2007 Practical Guide to Controlled Experiments on the Web is one line, at 95% confidence and 90% power. Run it twice, for the whole site and for the form the fix was about:

n = (4 × variants × std dev / difference)²

  • 111,957 visitors to see a sitewide move from 1.6% to 1.9%, so 56,000 per period
  • 1,536 people to see form completion go from 40% to 50%, so 768 per period
  • 14 months of before data at 4,000 visitors a month, against 1.3 months at the step

Same fix, same site, same confidence. One is unreadable for two years, the other inside a quarter.

What to set up before you change the page

Four things, none of them an afternoon's work. All four have to exist before you ship, because a baseline and a shipping date cannot be reconstructed from memory six weeks later, and that is the point at which most measurement quietly dies.

  1. The change list. One row per change, with its date and the one metric it should move. In Revslip that is the to-do list, and each row is tracked through to its revenue impact.
  2. The date marker. Google Analytics has built-in annotations, 1,000 per property with a 60 character title, shown on every report with a line graph. Mark the ship date.
  3. The step metric. Form starts to form submits. Cart to checkout. Trial to paid. Never the sitewide rate.
  4. Your own baseline noise. Eight weeks of that step metric, week by week, before you touch anything. The spread tells you what counts as movement here, which no benchmark can.

Want the numbers rather than the arithmetic? Run a free audit.


How to read the result four weeks later

Match the periods, same weekdays and same campaigns, then say the result out loud with its caveats attached. Kohavi's 2007 guide recommends whole weeks so day-of-week effects wash out, which applies to a before-and-after exactly as it does to a test.

The honest sentence: form completion went from 41% to 49% across four matched weeks, we changed two things on 12 March, traffic mix held steady. Weaker than a p-value, more useful than a made-up one.

What we cannot separate for you Revslip does not split your traffic, so it never runs a randomised test and never hands you a p-value. It cannot tell your fix apart from your market. What it does is narrow the question: every change is dated and tied to one metric, so when the number moves you have two candidates, not twenty.

When this is not worth measuring

When you can run a real test, run it. With a testing tool and enough traffic to reach power on your metric, a randomised split beats everything above and Revslip is the wrong instrument. Our piece on testing below the traffic threshold shows where that line sits.

Skip it for changes too small to carry a number. A footer link is not worth eight weeks of baseline. Spend the measurement on the three findings that arrived with euro figures attached.

One limit on our own numbers: the 134 businesses Revslip has audited are a biased sample, because a site gets audited when its owner already suspects something. Our figures describe suspicious sites, not all sites.

Sources: pre-experimental designs from the open-access VIVA research methods textbook, annotation limits from Google Analytics Help. Neither sells conversion tooling. We do, so read our arithmetic twice.

Questions people ask about measuring a fix

Three come up, and all three are one worry in different clothes: how long until I am allowed to believe the number moved because of something I did, rather than because March is simply a better month than February was?

Should I just run the audit again after the fix?

Run it, but treat it as a checklist rather than a measurement. A second audit confirms the defect is gone from the page. It cannot tell you what removing it was worth, because the page is not where money gets counted.

How long should I wait before reading the numbers?

Whole weeks, and enough of them to clear the sample size for your step metric. For a form reached by 600 people a month, about six weeks. For a sitewide rate on that traffic, longer than the fix will stay untouched.

What if I changed five things at the same time?

Then you have one measurement, not five, and crediting your favourite change is the mistake. Keep the combined figure, mark all five on the same date, and separate them next round by shipping one at a time.

Is this on your site?

Paste a URL and find out.

https://
No credit cardResults in 60 secondsWorks on any live site