Here is the setup almost every account runs at some point. Two ad sets in one campaign, the same audience, a different creative in each, and whichever one performs better gets the budget. It looks like a clean test. It is not a test at all.
What Meta actually does
The moment both ad sets go live, Meta does not split your audience into two even, isolated halves and let them run a fair race. Meta's own documentation is explicit: when ad sets run together like this, the system treats them in combination and skews delivery and budget toward whichever one performs better early. Within a few days it has effectively picked a front-runner and started starving the other.
So the ad set you labelled the "loser" never got a fair shot. It was throttled before it had the impressions to prove anything. You did not learn which creative is better. You learned which creative Meta happened to favour first, on a sample far too small to trust. The same rate-on-low-volume problem that fools people in reports is baked straight into the test.
Campaign budget optimization makes it worse
Turn on campaign budget optimization and the contamination deepens. Now the shared budget is actively redistributed toward the better-performing ad set, by design. The exact thing that makes CBO good for scaling a proven winner makes it useless for measuring an unproven one: the budget flows to the front-runner and away from the ad set you were trying to measure, before either has earned the verdict.
What a real test looks like
A valid A/B test does two things this setup cannot: it isolates a single variable, and it gives each version an even, uncontaminated audience. On Meta, the clean way to get both is the built-in A/B Test tool.
- It splits the audience randomly and without overlap. No one is shown both versions, so the groups stay statistically comparable and you avoid the audience-overlap problems that plague parallel ad sets.
- It holds everything constant except one variable. Creative, or audience, or placement, one at a time. Change two things and you cannot attribute the result.
- It runs long enough to mean something. Give each version two to three weeks and enough conversions to clear noise, not a couple of days. A test that ends before it leaves the learning phase is measuring instability.
If you insist on testing manually, you can, but then the versions have to be genuinely isolated, one change at a time, on separated audiences, and read only once the volume is real. Two ad sets sharing a campaign and an audience is none of those things.
The takeaways
- Two ad sets in one campaign share and skew delivery; Meta favours an early front-runner instead of splitting evenly.
- The "loser" is throttled before it gathers a fair sample, so the comparison is contaminated from day one.
- Campaign budget optimization worsens it by actively moving budget to the front-runner.
- Use Meta's A/B Test tool for a random, non-overlapping split on one variable, and run it two to three weeks.