We spent $52 on a video ad campaign and got an answer to a question we never asked. The question we did ask went completely unanswered. This is a write-up of how that happens, because the mechanism is invisible from every dashboard the platform shows you, and it is quietly wrecking more experiments than ours.
The setup
We publish documentary-style videos about corporate AI stories. We wanted to know which audience actually holds attention on the format: viewers in India, where this genre has an enormous home audience, or viewers in the US, UK, Canada and Australia, where our customers are.
So we built what looked like a clean head-to-head: one campaign, one creative, two ad groups. One targeted India, the other targeted the four tier-1 English markets. Same budget, $15 a day, and a pre-registered verdict: after 72 hours, compare average view duration between the arms, scale the winner or kill the campaign.
What actually happened
Over three days the campaign served 82,398 impressions to India and 4 to the entire tier-1 group.
Not four percent. Four impressions. Effectively 100% of the $52 went to one arm. Day-one analytics: 1,764 views, every single one from India, averaging 59 seconds of a seven-minute video. Of all the paid impressions, 2% reached the halfway mark.
The experiment compared nothing. There was no tier-1 arm. There was a cheap-geo buy wearing an experiment's clothes.
The mechanism
Nothing malfunctioned. That is the uncomfortable part.
Demand Gen campaigns optimize for engagement per dollar within the campaign's budget. A view from India costs a small fraction of a view from the US. Given one budget spanning both, every dollar sent to the expensive arm looks like waste to the objective function. The optimizer did its job with perfect fidelity. Its job just wasn't our question.
A shared budget across test arms is not a test. It is an auction the cheapest arm always wins.
And no surface warns you. The campaign dashboard showed healthy delivery, views arriving, spend on pace. Every top-level signal said working. The only place the truth lived was a geographic segment report that nothing ever prompts you to open.
Why we take this personally
This channel has been here before. Its early growth came from paid distribution aimed at cheap markets, and the result was a big audience number that produced no comments, no customers, and no truth. We turned that off and started over. Watching the same gravity reassert itself inside a $52 experiment was a useful reminder: money flows toward what is cheap to buy, not what is worth buying, and it does so through any structural gap you leave open.
The structure that prevents it
These are now standing rules for every paid campaign we run, ours or a client's:
1. One cost tier per campaign. If two audiences differ meaningfully in auction price, they never share a budget. Separate campaigns, separate caps.
2. The only shared component is the creative. Comparisons happen in the analysis, not inside a single campaign's allocator.
3. Gates are pre-registered. Threshold, metric, and kill date written down before the first dollar. Ours is 72 hours, scale or kill, judged on retention, never on view counts.
4. Verdicts come from segment reports. The campaign summary is marketing. The breakdown by geography, device, and placement is the data.
The reusable idea
An optimizer optimizes its objective, exactly, and nothing else. If your experimental design leaves it a degree of freedom, it will spend that freedom in whatever direction its objective points, and it will do so silently. You cannot brief an auction algorithm on your intentions. Constraints are the only language it reads, which means the structure of your campaign is the experiment. Design the structure, or the optimizer designs it for you.
Systems like this one, built for your business.
We build production AI systems for service businesses, and write up how they actually work, including the failures.
Book a free strategy call →