Tuesday, March 9, 2010

How To Build a Much More Useful Split-Tester in Excel Than Google's Website Optimizer

A Better Split Tester

Done in Excel Than

Google Website Optimizer

Google AdWords’ Website Optimizer is a great tool to run split-tests on landing pages in your AdWords account. You can, however, easily create a much more versatile split-tester with Excel that produces exactly the same result as the Website Optimizer but is much simpler to use.



This article and the video below will provide specific instructions on how to produce a split-tester in Excel that produces exactly the same result as the Website Optimizer because it runs the same statistical test (a one-tailed, two-sample, unpaired hypothesis test of proportion), but the Excel Split-Tester can be used in almost any marketing situation that would employ split-testing in any general way.

The Optimizer is a tool built into Google AdWords and can only be used in that medium.




Step-By-Step Video About How To Create an Excel Split-Tester That Is Much More Useful Than Google's Web Site Optimizer
(Is Your Sound Turned On?)



Google Web Site Optimizer in Action

Click On Image To See Enlarged View

Close-Up of the Above Screen

Click On Image To See Enlarged View

Excel Split-Tester Performing the Same Comparison


Click On Image To See Enlarged View

Click On Image To See Enlarged View
This is also the same Statistical test that the Excel Split-Tester performs


Example of Everyday Marketing Use of the Excel Split-Tester

The Excel split-tester that you’ll make can be applied to almost any marketing situation. For example, you could easily use your Excel split-tester to determine whether a change made to a direct mail campaign really made a difference.

The Google Website Optimizer runs the statistical hypothesis test on the number of clicks and number of conversion that each landing page elicits. The hypothesis test calculates the probability that the result obtained – one landing page having a higher conversion rate than the other – is true and not just the result of random luck.



How I Use the Excel Split-Tester

I use this Excel split-tester all the time in my job as an Internet marketing manager and I really enjoy its ease of use. All I have to do is plug a couple numbers right out of my AdWords account into the Excel split-tester and I have my answer in a second. The Excel split-tester has none of the set-up requirements that there are inherent with the Google AdWords Website Optimizer.

When I use the split-tester to test AdWords landing pages against each other, I will normally conclude that one landing page converts better when the split-tester states there is at least an 80% chance that the result obtained - one conversion rate is higher than the other – is a real and not just a chance occurrence.

The above video also provides a statistical derivation of the functionality of the split-tester.



Conclusion - It's the Greatest Thing Since Sliced Bread for the Marketer

If you are a marketer, you will get a lot of great use out of this extremely versatile and powerful tool.


Feel free to provide any comments to this article. Also, if you have any other ways of using Excel to optimize your AdWords account, let us know. Your input and opinions are highly valued!



If You Like This, Then Share It...
Dig this Stumble upon Delicious Technorati Reddit Buzz it Twitthis



Click Here to Download the Excel spreadsheet with the fully functioning Split-Tester shown here for only $7.

Also get the FREE BONUS

100+ page Excel Statistical Distribution Graphing eManual. This eManual shows you step-by-step how to create user-interactive graphs in Excel of all major statistical distributions.

Excel Master Series Blog Directory

Statistical Topics and Articles In Each Topic

Using Dummy Independent Variable Regression in Excel in 7 Steps To Perform Basic Conjoint Analysis

Using Dummy

Independent Variable

Regression in Excel in 7

Steps To Perform Basic

Conjoint Analysis

Overview of Dummy Independent Variable Regression

Dummy independent variable regression is technique that allows linear regression to be performed when one or more of the input independent variables are categorical. Categorical variables cannot act as the input independent variables in a linear regression analysis is their current form as nominal variables. Nominal variables are simply categorical labels that provide no indication of relative value or importance.

The categorical variables can be used as inputs to a linear regression analysis if each categorical variable is converted dummy variables that are binary, i.e., can only take the value of either 1 or 0. The number of binary variables for each choice category will equal the number of choices available for that category.

One dummy variable from each choice category must be discarded as an input for the linear regression analysis. The values of independent variables of a regression should not be predictable based upon the values of other independent variables. Any error called multicollinearity occurs if the values any independent variables can be predicted from the values of any other independent variables.

If one level of each attribute is removed it is not possible to predict the values of the remaining dummy variables of each attribute. It does not matter which dummy variable from each choice category is removed. Removing one level of each attribute does not affect the accuracy of the regression analysis, as will be demonstrated at the end of this article.

The independent variables is a linear regression analysis can be both binary dummy variables and continuous variables. The number of choices for each category should be relatively few or the regression analysis will quickly become unmanageably large as a result of the large number of dummy variables that would be needed for a large number of choices for categories.

Dummy Dependent Variables

Linear regression can be performed if the independent variables are categorical by applying the dummy variable conversion described in this article. Linear regression cannot be performed if the dependent (Y) variable is categorical.

The simplest case of a categorical dependent variable is a binary dependent variable. An example might be an attempt to use independent variables to predict the outcome of a binary event, such as a potential customer making a purchase or not. The technique to be applied in this circumstance is called Binary Logistic Regression. Here is a link to a series of articles in this blog which explain how this technique can be performed in Excel:

http://blog.excelmasterseries.com/2014/06/logistic-regression-overview.html

Overview of Conjoint Marketing Analysis

Conjoint analysis is a statistical technique employed by market research to create an equation that can be used to predict the degree of preference that people have for different combinations of product attributes. Conjoint analysis also enables market researchers to determine the relative level of importance that consumers on attribute choice categories and on the individual choices available in each category.

A product can be described by the attribute choices available to the consumer. At its most basic level conjoint analysis requires that a test subject assign a preference rating to each of all of the possible combinations of attribute choices available for a product. The preference rating scale goes from 1 (lowest preference) to 10 (highest preference).

The information obtained from this consumer test can be directly analyzed with linear regression if the categorical choices are converted to binary dummy variables. The resulting binary dummy variables can be used part of the set of input independent variables.

The output of this linear regression analysis is a regression equation that can be used to predict the test respondent’s preference rating for any combination of attribute choices. The coefficients of the regression equation indicate the relative degree of importance that the test respondent places on each of the attribute choices.

The following describes the 7-step process of using dummy independent variable regression to perform a very basic Conjoint analysis:

Step 1 – List All Attributes

List all of the available choices that a consumer has for one product. Starts by listing all of the overall attribute categories. In this case the attribute categories are brand, color, and price. Lists all of the available choices within each attribute category as follows:

Dummy Variable Regression in Excel - Attribute List

 

Step 2 – List All Possible Combinations of Attributes

Every possible combination of attributes should be listed. In actual Conjoint Analysis each unique combination of attributed is place on a separate card.

Dummy Variable Regression in Excel - Combination List

 

Step 3 – Rate All Combinations

The test subject will then rate each combination on a scale of preference from1 to 10 with 10 being the most desirable. Placing each unique combination on a separate card facilitates the rating process.

Dummy Variable Regression in Excel - Rating Combinations

 

Step 4 – Create Dummy Variables

In this step the categorical variables are converted to binary variables that can now as inputs to a linear regression analysis. Each level of each attribute will have its own binary dummy variable as shown below. The number of binary dummy variables for each attribute category will equal the number of choices available for that category. For example, there are three choices of brands with each choice being assigned to a single, binary dummy variable.

One dummy variable from each attribute category should be removed from the analysis. The values of independent variables of a regression should not be predictable based upon the values of other independent variables. Any error called multicollinearity occurs if the values any independent variables can be predicted from the values of any other independent variables.

If one level of each attribute is removed it is not possible to predict the values of the remaining dummy variables of each attribute. It does not matter which dummy variable from each choice category is removed. Removing one level of each attribute does not affect the accuracy of the regression analysis, as will be demonstrated at the end of this article.

The following are the listing of binary dummy variables for each of the attribute choice categories.

Dummy Variable Regression in Excel - Brand Dummy Variables

Dummy Variable Regression in Excel - Color Dummy Variables

Dummy Variable Regression in Excel - Price Dummy Variables

 

Step 5 – Arrange Data For Regression Analysis

The remaining dummy variables are input into the regression analysis as the independent variables while the preference rating is input as the dependent variable. Each record of data includes the binary dummy variables and preference rating from one of the cards. The data is arranged as follows:

Dummy Variable Regression in Excel - Regression Variables

 

Step 6 – Perform Regression in Excel

The Excel Regression dialogue box is then completed as follows:

Dummy Variable Regression in Excel - Completed Regression Dialogue Box
(Click On Image To See a Larger Version)

 

Step 7 – Analyze Regression Output

The Excel regression output appears as follows:

Dummy Variable Regression in Excel - Regression Output
(Click On Image To See a Larger Version)

 

The most important parts of the output are highlighted in the output and described as follows:

  1. The regression equation is calculated to be the following:

    Preference Rating = 5.61 + 1.67*(Brand B) + 3.5*(Brand C) + 1.33*(Blue) – 2.17*($100) – 4.17*($150)

    The value of each of the dummy variables is either 1 or 0 from the input data for each data record.


  2. The relatively high R Square, 0.87, indicates that the regression equation is a good predictor of Preference Rating. Approximately 87 percent of the variance of the Preference Rating is explained by the input variables.


  3. The low Significance of F (which is a p Value) indicates that the overall regression equation is significant with a high degree of validity.


  4. The low p Value for the Intercept and coefficients indicates that is significant with a high degree of validity.


Confirming the Validity of the Dummy Variable Regression Analysis Step

Plugging the values of the input independent variables for each data record creates the following comparison between the actual Preference Ratings given by the test subject and the Predicted Preference Ratings using the regression equation. The dummy variable regression analysis is seen to be relatively accurate. The removal of one dummy variable for each attribute choice category did not adversely affect the accuracy of the analysis.

The effect of removing a single dummy variable for each attribute choice category was to simply assign the value of 0 to coefficient that would be represented that dummy variable in the overall regression equation. The other coefficients have values relative to that value of 0.

The regression equation is shown by the Excel regression output to be the following:

Preference Rating = 5.61 + 1.67*(Brand B) + 3.5*(Brand C) + 1.33*(Blue) – 2.17*($100) – 4.17*($150)

If the dummy variables that were removed from the analysis would added back to the regression equation, the resulting equation would be the following:

Preference Rating = 5.61 + 0*(Brand A) + 1.67*(Brand B) + 3.5*(Brand C) + 0*(Red) + 1.33*(Blue) + 0*($50) – 2.17*($100) – 4.17*($150)

Both of the above regression equations would produce the same calculation of predicted Preference Rating.

The following image calculates the difference between the test respondent’s actual preference ratings for each combination and the preference ratings predicted by the regression equation.

Dummy Variable Regression in Excel - Actual and Predicted Preference Ratings
(Click On Image To See a Larger Version)

Dummy Variable Regression in Excel - Regression Equation in Excel
(Click On Image To See a Larger Version)

Dummy Variable Regression in Excel - Difference Between Actual and Predicted Preference Ratings

 

Excel Master Series Blog Directory

Statistical Topics and Articles In Each Topic