Six Months of Testing Contextual Intelligence, and Why We Stopped Trusting Happy Clients
Why do so many people in this industry still not believe contextual intelligence moves the needle on publisher revenue?
We've lost deals over exactly this. Prospects who heard the pitch and just didn't buy it. Some said outright they didn't think bid enrichment could actually lift revenue. Others put it on the market itself: not ready yet, too many DSPs still can't read these signals properly. Honestly, that second part isn't wrong across the board, some DSPs genuinely aren't set up to make full use of it yet. But not being universal isn't the same as not working. Where the signal does get read, and read well, we now have six months of tests showing the uplift is real.
For a long time, what we had to point to instead was clients telling us they were happy. Revenue felt up, things seemed to be working. But ask any of them for the actual number, the uplift, the fill rate change, the eCPM shift, and it turned into a shrug. "Better" isn't a number. And that was only half the story. We also had partners who tried it and got nothing back, no uplift, nothing changed. We didn't talk about that side nearly as much as the happy clients.
So six months ago we made a call: stop guessing, start testing properly. Not one page, not one partner, but test after test, across different publishers, different markets, different SSPs. User-level splits, minimum volume before we'd even look at a number, weeks of read time instead of a glance on day two. The kind of rigor you'd want if someone else was trying to sell this to you.
The Part Nobody Explains Well: Why This Actually Works
Here's the piece that gets glossed over in most contextual pitches. It's not really about pricing. Everyone assumes the win is buyers paying more for a well-labeled page, "sports," "finance," whatever the category. That effect is real, but small. Buyers do pay a bit more for relevance.
The bigger effect is something else entirely: volume. Without a real semantic read on the page, a huge amount of completely normal content never gets bought at all. Think of a keyword filter scanning a cooking article that mentions "knife skills" the same way it would flag an assault report, no understanding that one is about dicing an onion and the other is a public safety concern. That's the failure mode: shallow keyword matching can't tell the difference, so it blocks or discounts content that was never risky to begin with. That inventory isn't underpriced. It's unsold. Once there's a signal built on real content understanding, not just string matching, that inventory becomes sellable in the first place. That's the mechanism doing most of the work: impressions that used to earn nothing finally getting bought.
And it doesn't stop at existing buyers paying more. In one of our clearest tests, 17 buyer networks that had never bid on the inventory before showed up once the content became readable to them. Not better prices from familiar names. New demand that simply couldn't see the inventory until it could be classified properly. We can also see, at the buyer level, which types of demand respond fastest to this, and it's not random, there's a clear pattern by DSP and trading desk. That's a conversation worth having directly rather than a stat worth generalizing here.
That's the part most pitches skip, and it's also the part that's easiest to get wrong if your testing methodology isn't tight.
If you've ever run a contextual pilot that underperformed, or you're skeptical this holds up outside a slide deck, stick around. Next week I'm walking through the actual test-by-test numbers, the ones that worked, the ones that flatly didn't, and the mistakes it took to get results we now trust.