Split/MV Testing and Statistical Significance

Alejandro

The HOTH - Google Me.
May 21, 2009
465
5
0
Chi-city
My Story (aka. "Why I'm Mildly Retarded for not Split Testing Earlier" - feel free to skip)

So back in September I started a a site for an info product. My campaign was solid. I dominated the entire organic keyword niche for the product, and for a while, shit was gravy.

I quit my job. Went to europe for 5 weeks. Came back. Still making money. Dope.

Soon after, sales started declining. What happened? Well, it looks like competition for the product nearly DOUBLED. Damn.

What else?

Having been at #1 for so long, I was easy picking for someone to rip off my LP once all the new kids came to the party.

Now, I consider myself a pretty smart person, but jesus christ did it take me a long time to realize that MAYBE all of those LP rip offs who were actually buying PPC ads, and therefore ranking above me, were making my ORIGINAL LP look like BULLSHIT. It totally killed the credibility of my flog author (ironic statement, i know, haha).

Lightbulb.

So last friday, at 3am, I found a completely different theme for my landing page (Wordpress based) that looked NOTHING like my previous LP or any of its rip offs. I spent a couple hours revamping the entire thing, and bam, my first "a/b" test if you will.

My "Holy Shit Split Testing is Awesome!!!" Moment
(aka. "Why I'm STILL Retarded for not Split Testing Earlier" - feel free to skip)


So everyday, for the past 5/6 days, I've been looking at my analytics. Have bounce rates changed? Nope, not really. How about time on page? Pretty much the same. CTR? Still roughly the same.

Damn. I almost wouldn't have even minded if my stats when DOWN, I was just expecting SOME kind of change from a 100% aesthetic overhaul.

Like I said, my product is a CB product, and its been a bitch to set up some sort of actual conversion tracking, so for the most part, I'd pay attention to CTR and pray that conversion rate stayed roughly the same.

Nuh-fucking-uh.

What I found almost made me shit my pants.

Earnings per visit to site (after returns) From $.22 to $.34 - 55% increase!
Earnings per visit (before returns) - From $.235 to $.38 - 61.7% Increase!
Conversion Rate (#Sales/Visit to LP) - From .9% to 1.4% - 55% Increase!
^^^ (before I bet flamed for such a low CR, let me say a) lick my balls b) this is CPS not, cpa, so for a page that actually got people to willingly spend $40 without any split testing, I think i did alright)

Holy fucking balls. But, like I mentioned previously, I generally consider myself "smart," so I realize there's a decent chance my findings weren't statistically significant. This evening, I inputed my numbers into Split Test A/B Test Marketing Calculator - landing page-ppc-email split tests | SplitTestCalculator.com (awesome tool! recommended!) and, to my dismay, found that my morning numbers were only 46% significant, haha. By evening, after a slow day, they were only 16% significant. FML.

Now I've Caught the Split Test Bug - Please Help Me (DON'T SKIP!)

After realizing the power of having statistically significant information, after 6 months of living off internet profits, it FINALLY hit me how valuable paid traffic is for optimizing pages. I immediately called up a friend who is the head ppc-guy for a top e-commerce company and we immediately decided to pursue a joint venture.

His job will be to drive ppc traffic. My job will be to split test like crazy and eventually get free organic traffic for our top converting terms.

MY QUESTIONS

  1. Theoretically, if the difference in conversion rates is small enough, it could take ages to achieve statistical significance. What rule of thumb do you use for dropping a test when it looks like it may take weeks to get enough data? Is there a certain amount of views at which you cut it off? A certain amount of conversions?
  2. Let's say you do decide to end the test. One of the 2 variants you were testing had done better, but you were only able to achieve 50% significance before you reached the threshold. Would you then keep your page the same, or would you change the page to reflect the variant that performed better, even if the difference was statistically insignificant?
  3. Any other tools/resources you can offer on this subject, I am all ears. I think its pretty fascinating.

Again, for all others new to testing, Split Test A/B Test Marketing Calculator - landing page-ppc-email split tests | SplitTestCalculator.com is a great tool that will keep you from blowing your load too early like I did this morning :)

Thank in advance for any answers.

P.S. I think its od that for the billions of times people tell noobs "test, test, test," very few actually discuss the methodology on here. Not hating, just observing. Holla.
 


<3 MVT - I'm also kind of obsessing over the methodology myself lately.

Here's a nice MVT tool to help in your endeavor: Visual Website Optimizer - World's easiest to use A/B, Split and Multivariate testing tool - I think it's still in beta, but you can find an invite code on their blog. It's pretty sweet.

I would also suggest http://www.clicktale.com to help you see where people are dropping off in the sales funnel or are losing interest and leaving. It's pretty helpful at giving you qualitative data. It's nice to be able to see things such as maybe a piece of content half way down your page getting a ton of attention that you would be best off having towards the top to hook more people. http://www.openwebanalytics.com is an open source alternative that I recently came across that has the same sort of dom-recording that clicktale does, only far less features so use it at your own risk.

As far as your questions. The first thing you need to do is get -real- conversion data and not just "guesstimations". Since you obviously have a well performing product, you may want to consider switching to a billing solution that will allow you to properly do this if you can't with clickbank. It may seem like a pain in the ass at first, but real conversion data is somewhat critical to your success.

I'm not a statistician, so I can't help you too much with the statistical significance stuff other then say I would operate off of views and optimize for either eCPM if you're doing display or EPC if you just stick with PPC. The reason you want to focus on either of those two numbers is because you may find that one ad may bring you in more clicks over another, but the one with less clicks actually has more conversions for whatever reason.

Your answer to #2 would be dependent on how much better the winner did. I would go ahead and say if you have to ask this question, chances are you should be running the test longer for a more definitive answer. Just my 2c, GL!
 
Did you run half your traffic to each page? Or did you switch to the new page and compare against previous numbers? A/B significance testing is geared to the first situation, not the second.

#1 - Depending on traffic volumes, you're going to get an expectation of how long it takes. For example, where I work we normally run a split test for a week, but we know within a day if things are going well, and within a couple of days are already working on the next iteration.

#2 - If it's not significant, it's not different ;) In that case I'd probably look closer at patterns in the traffic, or try something different.

#3 - I've been using R (The R Project for Statistical Computing) to try and identify patterns in the data and do the calculations. The books "Head First Statistics" and "Head First Data Analysis" are good for people like me who are a few years out of practice with the math.

My split testing epiphany - I made some changes to a certain algorithm we use to predict upsells and ran a split test, which showed no difference. Spent some time analyzing data and found a distinct pattern: the new algo worked on one segment of traffic but the opposite on another. Rewrote the algo to only do its magic on the segment where we were doing well. The next a/b test showed it performed awesomely, after that test finished we threw all the traffic at the better performing algo and boom... I paid for myself a couple of times over.

Sean
 
MY QUESTIONS

  1. Theoretically, if the difference in conversion rates is small enough, it could take ages to achieve statistical significance. What rule of thumb do you use for dropping a test when it looks like it may take weeks to get enough data? Is there a certain amount of views at which you cut it off? A certain amount of conversions?
  2. Let's say you do decide to end the test. One of the 2 variants you were testing had done better, but you were only able to achieve 50% significance before you reached the threshold. Would you then keep your page the same, or would you change the page to reflect the variant that performed better, even if the difference was statistically insignificant?
  3. Any other tools/resources you can offer on this subject, I am all ears. I think its pretty fascinating.
To answer question #1:

It depends on the conversion rate but at 1-2% I'll end the test after 50 sales (combined) and stick with the control if there's not a wide gap.

To answer question #2:

50% significance isn't significant at all. It's a coin toss. Keep the test going until you hit sales numbers above. If it still stays at 50% or close to it, end test and start a new one.

I've had tests at 90% significance (> 100 sales) turn out not to be an improvement once I ended the test and made the winner the new control.

To answer question #3:

Use Split Test Accelerator and focus on Taguchi instead of split testing. You can test multiple areas of your page at once to zero in on things that are most likely to significantly impact response. Typically this will be the headline and elements above the fold and the offer but sometimes you'll be surprised at the little things that can affect response.

Once you find the areas that have an impact on your response, THEN start split testing that to improve.

And since you're working with CB here's another tip. The look and feel of your order form plays a big role on your conversions. Now, you can't change the look and feel of CB's order form but you can definitely make your site's color scheme match CB's. It's something you should definitely test. If that fails to increase response, then put a video on your page walking people through the checkout process (just blur out the name, address, and CC fields for obvious reasons) and show them the download page. That way they'll be more reassured that you're not going to take their CC number and run without delivering the product.
 
I haven't had a chance to go through everything suggested here so far, but I nominate this for a sticky thread already.
 
I just wanted to raise something i read recently, i forget where, but it was on some Google Certified Consultants blog (i think) regarding statistical significance.

He said something which i'd never truly thought of properly before. He says something we all do is assume visitors are quantifiable; that all visitors are the same when considered on a per thousand or some similar measure - we generalise them.

He said that, and i know this is super-obvious but bare with me, every single person is different. Therefore every single thing that you can test changes with each person is going to give a different outcome, each time. He then qualifies what he says (and kind of makes more sense) by explaining that he does split tests that i've never heard anyone else talk of before.

What he does is conduct split tests with nothing changing. Same ads, same words, same lps, the whole lot. He would do this across thousands of vistors and says that he would have a natural variance of between 1%-3% regardless.

Just thought i'd share that. :)
 
What he does is conduct split tests with nothing changing. Same ads, same words, same lps, the whole lot. He would do this across thousands of vistors and says that he would have a natural variance of between 1%-3% regardless.

That doesn't make any sense. How can you conduct a test without changing anything? Furthermore, if you didn't change anything how can you have any variance at all? Did you omit some important detail?
 
That doesn't make any sense. How can you conduct a test without changing anything? Furthermore, if you didn't change anything how can you have any variance at all? Did you omit some important detail?

Are you serious?

Actually i think i can see what you're on about, you just need to be able to splt the traffic and be able to see conversions. For instance, set up 2 identical ads in an ad group on adwords with everything being identical, run them evenly, track conversions.
 
Are you serious?

Actually i think i can see what you're on about, you just need to be able to splt the traffic and be able to see conversions. For instance, set up 2 identical ads in an ad group on adwords with everything being identical, run them evenly, track conversions.

It's called an A/A test, and the purpose is to make sure your testing is correct. In the cases where you can't do a true split, this helps you make sure your results would be correct.

For example, if you decided that to test the hypothesis that "doubling my bids on tuesdays will make me more money", then you'd have to set up a test and control keyword group. An A/A test should show that your metric is similar, otherwise you've messed up your groups or the test itself.
What he does is conduct split tests with nothing changing. Same ads, same words, same lps, the whole lot. He would do this across thousands of vistors and says that he would have a natural variance of between 1%-3% regardless.

The idea behind the confidence of the test is to point out those 1-3% of differences are there by random chance. Yes, if you send 10000 clicks equally split across 2 LPs then the output will differ. However, the confidence interval tells you if that was to be expected.

Practical Guide to Controlled Experiments on the Web: Listen to Your Customers not to the HiPPO is a great paper on split testing and running proper experiments.

Sean
 
  • Like
Reactions: uguxseo
Are you serious?

Actually i think i can see what you're on about, you just need to be able to splt the traffic and be able to see conversions. For instance, set up 2 identical ads in an ad group on adwords with everything being identical, run them evenly, track conversions.

Oh OK... I see what you're referring to. Totally agree that there can a variance as high as 1-3% by doing that. That's why I mentioned above that even if you hit 90% confidence interval it still might be a fluke.

However, if you were to run several thousand clicks across 2 identical Adwords ads with same LP, KWs, etc. eventually the variance will decrease dramatically to the point of evening out.
 
It's called an A/A test, and the purpose is to make sure your testing is correct. In the cases where you can't do a true split, this helps you make sure your results would be correct.

For example, if you decided that to test the hypothesis that "doubling my bids on tuesdays will make me more money", then you'd have to set up a test and control keyword group. An A/A test should show that your metric is similar, otherwise you've messed up your groups or the test itself.


The idea behind the confidence of the test is to point out those 1-3% of differences are there by random chance. Yes, if you send 10000 clicks equally split across 2 LPs then the output will differ. However, the confidence interval tells you if that was to be expected.

Practical Guide to Controlled Experiments on the Web: Listen to Your Customers not to the HiPPO is a great paper on split testing and running proper experiments.

Sean

Right, agreed. One of things I do when running a Taguchi test is to leave one factor blank. That way, if you have factors in there that appear to be winners and the blank factors are more or less equal, you have one more
vote of confidence that your test is accurate.