What makes it a test, not a guess
A valid A/B test needs three things: exactly one variable changed, an even and non-overlapping audience split so no one sees both versions, and enough volume for the result to be stable rather than noise. Miss any one of them and the outcome is not attributable to the change you made.
The mistake that isn't a test
On Meta, running two ad sets in one campaign feels like a split test but is not one. The system does not split the audience evenly; it skews delivery and budget toward the early front-runner, and campaign budget optimization makes that worse by actively moving budget to the leader. The clean method is the platform's built-in A/B Test tool, which enforces the random, non-overlapping split. We work through it in two ad sets in one campaign is not an A/B test. And whatever the split, judge it on volume, not an early conversion rate, and wait out the learning period before you read the result.