Almost Every AI Visibility Case Study You Have Read Is Superstition

#SEO#AI


Almost every AI search case study you have read is superstition. The people writing them are not dim or dishonest. They are pointing at a wildly noisy number and telling a clean causal story about it. Somebody changed their content, their citation count went up, and now there is a case study. The trouble is that the number moves on its own, constantly, whether or not anyone touches the site.

The Retrieval Layer Is Constantly Shifting

AI answers rely on a retrieval layer that pulls information from the web before the model writes a response, and it does not sit still. As David Konitzny showed this week in his research on invisible shifts, query fan-outs change constantly, the sources a model prefers get swapped without warning, and model updates happen quietly in the background.

Konitzny tracked citation and retrieval rates over time and watched them move while the target sites sat still. In May, ChatGPT quietly started appending “reddit” to its background queries, and Reddit’s share of retrieved sources jumped from under 2% to 5%. Reddit did nothing to earn that. Over the same stretch, Wikipedia’s share of reference retrievals fell from about 40% in March to 7% by July. Wikipedia did nothing to lose it. The model reweighted its own authority heuristics, and every brand riding on those sources moved along with it.

So if you rewrite your content on Monday and see a 40% jump in citations on Friday, you will naturally credit your work. Maybe it helped. Maybe the model just started fanning out its queries differently that week. From inside your own dashboard, those two stories look exactly the same.

We Trust The One Number We Can Count Far Too Much

Here is the part we are not honest about. We track citations because we can, and then we quietly promote a thing we can count into a thing we can trust. It is the old joke about the drunk searching for his keys under the streetlight because that is where the light is. The count is real. It is also one of the fuzziest measurements in marketing, and we keep reporting it as if two decimal places made it precise.

A control group is the cleanest way to cut through that noise. Hold one set of pages constant, change another, and see whether the difference survives. When you can run that, run it. This week a group of practitioners including Tory Gray, Shaun Davidson, and Anne Berlin published the AI/LLM Experiment Community Standard 0.1, a checklist for tests that actually hold up: hypotheses, controls, constants, and repeatability. It is a good bar to reach for.

But Tom, We Cannot Wait Around For A Perfect Experiment

You often cannot reach for it, and you should not pretend otherwise. When a technology is brand new and the ground is moving under you, intuition and speed genuinely beat waiting for a clean experiment. Some of the best calls I have made came from acting on a hunch long before there was any way to prove it. Chaos rewards the people who move.

So this is not a demand for a randomized trial before you are allowed to have an opinion. It is a demand for honesty about what you actually know. Move fast on your gut when the moment calls for it. Just call it what it is: a bet you are placing with incomplete information, not a result your dashboard proved. The moment you start treating a fuzzy citation count as hard evidence that your tactic worked, you have stopped doing marketing and started reading tea leaves.

Trust your intuition. Trust the number a lot less. Almost every AI search case study you have read is superstition, and the first step out of it is being honest about how little the measurement actually tells you.