Video A/B testing: how to split-test a sales video properly
Most video A/B tests are decided before they start because they count plays, not leads. The metric, the sample, the size of the bet, how to run a real split behind one link, and why the better question is which version wins for whom.
Most video A/B tests are decided before they start, because the person running them picked the wrong thing to count. They test two thumbnails and declare a winner on play rate. They test two hooks and declare a winner on watch time. Both numbers move, both feel like progress, and neither one is the reason the video exists. A sales video exists to produce people you can sell to. That is the metric, and everything below follows from refusing to test against anything else.
Why most video A/B tests fail?
Three ways, usually at once. The metric is upstream of money (plays, watch time, clicks), so the winner is the version that entertains, not the version that converts; we made that argument in views don’t buy. The sample is too small, so the test is called on noise. And the variable is too small: two thumbnails differ by a hair, the result is a coin flip, and the lesson is nothing. Fix the metric, the sample, and the size of the bet, in that order.
A test that counts plays finds the version people watch. A test that counts leads finds the version people buy from.
Pick the metric that pays
Count captured leads per viewer, and if you can, qualified leads per viewer: the ones who answered the qualifying question the right way or booked the call. Watch time and play rate are diagnostics, useful for explaining a result, useless for declaring one. If two versions produce the same number of leads and one has a higher watch time, you have learned something about attention and nothing about revenue.
Test one big thing, not two small ones
The tests that teach you something are the ones where the versions differ in a way you could explain to a colleague in a sentence: the hook (problem-first versus proof-first), the offer (call versus trial), the length (the nine-minute cut versus the four), the presenter. A thumbnail test is a fine third test. It is a terrible first one.
Run it behind one link
In StreamAgent the testing instrument is a Smart Rotation, and it is worth being precise about what it is, because it is not a coin flip between two variants. A rotation holds audiences (rules about who is arriving) that each point at a version, and a default for everyone else. A control share, ten percent unless you change it, always sees the default no matter what, so you have a baseline. Lift is how much better the routed viewers convert than that baseline.
To run a plain A/B with it: make version A the default, raise the control split to fifty percent, and add one audience whose only condition matches every visitor, pointing at version B. Half the traffic sees A as the baseline, half sees B as the routed arm, and the lift reading is the verdict. Every version serves from the same link, so the test never touches your ad URLs, and a preview link per lane lets you watch each version without polluting the numbers.
Let it run until the lift is real
The lift reading appears once both arms have thirty visitors, and thirty is the floor, not the target. Call a test early and you are reading weather. Decide in advance how many leads per arm you want to see before you believe the answer, and treat the first days of any test as the period in which nothing you see is true.
Then ask the better question
A/B asks which version wins for everyone. The more valuable question is which version wins for whom, and it is the one the rotation was actually built for. Paid clicks and organic visitors, first-time and returning, people who already watched the demo and people who have not, visitors who arrived from a particular campaign: each can be an audience with its own version, and the control share measures whether the whole routing arrangement beats the single-video baseline. When a campaign is quietly converting better in the default lane, the rotation suggests promoting it to its own audience. The rotation also runs on a calendar, so a launch-week version can have a start and an end date without anyone remembering to switch it off.
Promote the winner and keep the instrument
When a version wins, make it the default and leave the rotation in place, because the next test is always coming. Versions are cheap to make when the pitch is a route rather than a single file: swap the hook clip, keep everything downstream, test again.
Test the pitch, not the thumbnail. Count the leads, not the plays.
One platform, seven layers
This article covered one slice. The machine ships whole: record it, get it found, arm it, test it, rank the leads, train the ads, and ask your AI how it is going.
Common questions
- What is video A/B testing?
- Serving two or more versions of a video to comparable viewers and measuring which one produces more of the outcome you care about. For a sales video the outcome is captured or qualified leads per viewer, not plays or watch time, which are diagnostics rather than results.
- How much traffic do you need to A/B test a video?
- More than feels necessary. In StreamAgent the lift reading appears once both the routed and control arms have thirty visitors, and that is the floor, not the target. Decide in advance how many leads per arm you want to see before you believe the answer, and treat the early days as noise.
- How do you A/B test a video in StreamAgent?
- With a Smart Rotation. Make version A the default, raise the control split to fifty percent, and add one audience whose condition matches every visitor pointing at version B. Half the traffic sees A as the baseline, half sees B as the routed version, and the lift reading compares them on leads. Every version serves from the same link, with a preview link per version.
- What is the difference between A/B testing and audience routing?
- A/B asks which version wins for everyone. Audience routing asks which version wins for whom: paid clicks versus organic, new versus returning, people who watched the demo versus people who have not. Smart Rotations do both, holding out a control share so the whole arrangement is measured against a single-video baseline.