- Blog
- Web Design
- A Guide to A/B Testing for Dealer Groups
A Guide to A/B Testing for Dealer Groups
Author: Martin Dew
Published:
A Guide to A/B Testing for Dealer Groups
TL;DR
A/B testing helps dealer groups make better website decisions by comparing two versions of a page at the same time, rather than relying on opinion or before-and-after changes. Start with a clear customer-journey problem, test one meaningful change and measure an agreed primary outcome alongside guardrails such as form errors, abandonment and lead validation.
User testing and A/B testing work best together. User testing helps identify points of confusion and explain behaviour; A/B testing shows whether a defined change performs better with a live, comparable audience. Plan the sample size and stopping rule before the test begins, assess both statistical significance and the size of the improvement, and keep the experiment safe for search by avoiding cloaking, using canonical tags for alternate URLs and removing test components when the test ends.
A/B testing is a formal way of testing website changes by showing a portion of visitors the original version of a page (the A page) and another portion the B page. The aim isn’t to make a page look nicer or newer. It is to learn whether a specific change helps more visitors complete a useful next step.
For a dealer group, that can mean testing a clearer route to a test-drive request, a different service-booking prompt, or a more useful location-page call to action. The value comes from replacing assumptions with evidence by using a formal approach to experiments.
The approach is superior to “before and after” testing because it eliminates more external factors, such as seasonality or time of day, since some users see the A page and others the B page at the same time.
A good A/B test asks one focused question, measures one meaningful outcome, and changes only what the group is genuinely prepared to keep if it proves better.
Why A/B testing matters for dealer groups
A dealer-group website has more moving parts than a single-location site. It usually supports several brands, locations, service departments, stock feeds, and local customer journeys. A design or journey decision that works well for one page type may not be the best approach across the board.
Testing provides a way to make controlled, data-driven decisions about elements that you can control, such as page structure, content hierarchy and calls to action. It is not a way to test whether a dealership should discount a vehicle, source different stock or change its sales process. Those are operational decisions. The useful website question is narrower: does this version of the digital journey make it clearer and easier for a customer to take the next appropriate action?
User testing and A/B testing answer different questions.
User testing on your website and A/B testing are not alternatives but complementary. A/B testing is quantitative. It can show which live version generated a higher rate of an agreed action from a comparable audience. It rarely explains why people behaved differently.
User testing is qualitative. It can reveal whether a person understands a label, notices a next step, hesitates over a form, interprets an offer correctly or encounters a point of confusion. It is especially useful before a live experiment because it helps the team form a stronger hypothesis, and after an experiment because it can help interpret an unexpected result.
It’s also important to understand that an A/B test is not a preference test. In a preference exercise, people may see pages A and B, and say which they prefer. That can be useful design feedback, but it does not show how separate live audiences behave when each sees only one version in a real task.
When combining user testing and A/B testing, a typical sequence is: observe a problem through user research; form a hypothesis; A/B test one proposed solution; then use follow-up research where the result needs explanation. It’s during those research stages that user testing can be useful, alongside tools like heatmaps and analytics data.
What makes a useful test candidate?
When we work with dealer groups on Conversion Rate Optimisation initiatives and agree to a program of A/B testing, the first thing that usually happens is that there is a desire to test and improve a lot of things because the website is so large. So, we need to consider ideas and identify which areas are strong candidates for A/B testing and where we should use other methods.
For A/B testing, you should find a customer journey with an identifiable point of friction or uncertainty. These will usually present themselves when examining user journeys in Google Analytics or heat-mapping data. Be careful not to assume there is a problem based on a small number of recorded sessions - these can be outliers - you should make sure you have enough aggregated data to see a trend.
Here are some examples of typical candidates for A/B testing:
|
Website area |
Focused test question |
Meaningful primary measure |
|
Vehicle detail page |
Does a clearer, more prominent enquiry action increase completed enquiries without reducing the quality of the next step? |
Completed, valid enquiry event |
|
Test-drive route |
Does presenting availability or appointment information earlier encourage more completed test-drive requests? |
Completed test-drive request |
|
Service booking page |
Does a shorter explanation of the booking process produce more completed bookings? |
Completed bookings |
|
Location page |
Does a location-specific call to action make it easier for visitors to contact the right dealership? |
Location-specific call, form or direction action |
|
Part-exchange journey |
Does clearer preparation information increase the number of user valuation starts? |
Number of users starting a valuation |
|
Campaign landing page |
Does a simpler page hierarchy help visitors reach the campaign’s intended action? |
Defined campaign conversion event (e.g. ‘Book Your Place’, ‘Register Your Interest’) |
A useful A/B test is specific enough to be disproved. “Make the vehicle detail page better” is not a test hypothesis.
The six steps of a useful test
1. Define the decision before building the variation
You should define the question, the proposed change, the primary measure, the audience and the decision that you will make after the result. This prevents a test from becoming a vague design-preference exercise. For example:
|
Test element |
Example |
|
Question |
Can mobile visitors find the test-drive action more easily? |
|
Hypothesis |
Presenting the action after the key vehicle facts (make, model, variant, price, mileage), rather than below the full description, will increase completed test-drive requests. |
|
Audience |
Mobile visitors to selected vehicle detail pages. |
|
Primary measure |
Completed test-drive request event. |
|
Guardrails |
Form errors, abandonment, and lead validation where it is available. |
|
Decision |
Retain the alternative only if it improves test drive bookings without a material negative impact on the guardrails. |
2. Change one meaningful thing
If a B page changes several things at once, it may produce a result but will not explain why. Begin with a single, meaningful difference.
This does not mean every test should be a button-colour experiment. A meaningful single change can be a clearer page hierarchy, or a shorter form with the same essential data. It means the test should answer one question at a time.
3. Choose the right audience and protect comparability
A and B pages should run simultaneously and be assigned randomly within the agreed-upon audience.
You should decide in advance whether a test applies to one location, one franchise, one page template or the wider group. This helps to make the test valid. A good illustration is premium-brand and value-brand customer journeys. The journey to a test drive for these audiences is likely different, so combining those audiences in a single test could invalidate the results.
4. Define guardrails and measure the full journey
Guardrails are a crucial part of A/B testing. These measure the overall effect of switching from an A page to a B page, beyond just the test’s goal (e.g., test-drive bookings). As an extreme example, let’s say you removed all navigation and distractions from a vehicle detail page, leaving a prominent test drive button. You may end up with more test-drive bookings from the B page, but there will almost certainly be negative impacts due to the disrupted customer journeys in that version. Overall, your financial position may be worse using the B page - the guardrails make sure this is part of the equation.
5. Plan the sample size and stopping rule before launch
How long does it take to complete an A/B test? It depends on traffic and the frequency of the outcome being measured. There is no universal “run it for seven days” rule. A test involving a rare completed enquiry will need more exposure than one measuring a common navigation click. The outcome needs to be statistically significant based on mathematical tests, or you are basing decisions on random outcomes.
Remember: a result that doesn’t show a meaningful difference is still useful. It tells you not to spend time rolling out a change based solely on preference.
The mathematics: what statistical significance actually means
Statistical significance is important to understand because it is what A/B test outcomes are based on. on. An A/B test compares conversion rates, not opinions. If 100 of 2,000 visitors complete an action in variant A and 120 of 2,000 complete it in variant B, the observed conversion rates are:
A = 100 ÷ 2,000 = 5.0%
B = 120 ÷ 2,000 = 6.0%
Absolute difference = 6.0% − 5.0% = 1.0 percentage point
Relative lift = (6.0% − 5.0% ) ÷ 5.0% = 20%
The important question is whether the observed 1.0-point difference is likely to reflect a real underlying difference, rather than ordinary random variation in which visitors happened to convert more often in B during that period.
The null hypothesis and p-value
A regular A/B test has a null hypothesis and a p-value.
Null hypothesis: The A and B pages have the same conversion rate
p-value: Assuming the null hypothesis is true, how surprising is it to get the result you did, or a result where the conversions differ even more than the result observed.
If your A/B test reports p ≤ 0.05, the result is described as statistically significant at the 5% level. In plain English, the data would be relatively unusual if there were truly no difference between the variants.
That does not mean there is a 95% probability that B is better, nor does it prove that the result is commercially worthwhile. A p-value is evidence about compatibility with the no-difference hypothesis; it is not a measure of the size or value of an improvement. You should consider the size of the effect alongside statistical significance.
Effect size: is the change worth having?
For a dealer group, a statistically significant change can still be too small to justify development, training or group-wide rollout. Report both the absolute and relative difference.
|
Measure |
Example |
Why it matters |
|
Absolute difference |
5.0% to 6.0% = +1.0 percentage point |
Makes the scale of the change clear |
|
Relative lift |
(6.0% − 5.0%) ÷ 5.0% = +20% |
Helps compare performance against the starting rate |
|
Confidence interval |
A plausible range around the estimated difference |
Shows whether the data are compatible with a small, negligible or meaningful change |
|
Guardrail outcome |
No material increase in form errors or abandonment |
Helps ensure an apparent improvement did not damage another part of the journey |
Power, sample size and minimum detectable effect
When you’re thinking about test candidates, a common issue is making sure that there is enough data to get statistical significance within a sensible timescale. A test needs enough visitors to have a realistic chance of detecting an effect that is worth acting on. This is called statistical power. Power is the probability of detecting a real effect of the chosen size when it exists; 80% is a commonly used planning convention.
There are mathematical techniques to calculate statistical power, but we won’t go into that much detail here.
Keep A/B tests safe for search visibility.
A/B testing does not have to harm search performance, but it needs to be implemented honestly. You should not show Googlebot one version of a page and people another; that would be seen as cloaking.
Where a test uses separate URLs, use a rel="canonical" link on variation URLs that points to the original URL. If visitors are temporarily redirected to a variation URL, use a 302 temporary redirect, not a 301 permanent redirect. Finally, you should run tests only as long as necessary and ensure that test components are removed at the end of each test.
|
Test implementation |
Search-safe approach |
|
One URL, dynamic variation |
Assign variants consistently to real visitors and do not create a special version only for search crawlers. |
|
Separate variation URLs |
Canonicalise the alternatives to the original page. |
|
Temporary redirect to a variation |
Use a 302 redirect, not a permanent 301. |
|
Completed test |
Implement the agreed result and remove redundant scripts, markup and alternate pages promptly. |
A reasonable first test for most dealer groups
For a good first test, aim to start with a high-traffic page type where the proposed change is simple, customer-focused, and low-risk. An example could be a test drive or enquiry action on a vehicle detail page: retain the same approved information and destination, but test whether their placement and wording make the next step easier to understand.
If you would like help planning a safe, measurable experiment on a dealer-group website, the Autoweb team would love to hear from you.
