FIELD NOTES · AI SEARCH · 2026

THE BLACK BOX ALREADY PICKED YOUR SOURCES

On crawlability, citation, and the quiet consolidation of how we find things out.

I started with a hunch I couldn’t shake. Ask Gemini, ChatGPT, and Perplexity the same question and you get three answers that look like siblings: confident, formatted, sprinkled with citations. But they aren’t reading the same internet. One is reading a smaller version of it. One is reading a version shaped around you. And the part where each of them decided what to ignore is the part you never see.

I wanted to know if anyone was actually studying this, whether publishers understood what was happening to them, and whether the rest of us had even noticed. The research turned out to be further along than I expected, and worse.

The footnote is broken

Here’s the thing that should be simplest. An AI tool gives you an answer, attaches a source, and you assume the source backs up the claim. That’s the whole deal with a citation. Don’t trust me, go check.

In 2023 a Stanford team, Nelson Liu, Tianyi Zhang, and Percy Liang, did the tedious work of checking. They had humans audit four generative search engines line by line. On average, only about half the generated sentences were fully supported by their citations. Only about three quarters of the citations actually backed the sentence they hung off of.[1] Half the time, the confident sentence and its neat little link weren’t really connected.

Figure 1. Even when the cited page is right there, the link to the claim often isn’t.

But the finding I keep thinking about is one they almost tucked away. The answers people rated most useful had worse citations, not better. More fluent, more satisfying, less sourced. The system is good at producing the feeling of being informed, and that feeling runs opposite to whether the sourcing holds.

Then there’s a distinction I didn’t grasp at first and now think is the whole game. A 2024 paper, ‘Correctness is not Faithfulness in RAG Attributions,’ pulled apart two things most of us mash together.[2] A citation can be correct: the cited page really does contain the claim. And it can be faithful: the model actually used that page to get there. Those aren’t the same thing. In one model they tested, more than half the citations were what they called post-rationalized. The model answered from memory, then went looking for a page that matched, and stapled it on after the fact.

That’s the trick. The citation isn’t a record of where the answer came from. It’s produced afterward to justify an answer the model already had. It looks like provenance. It’s closer to an alibi.

The interpretability people are starting to confirm this inside the machine. A 2026 paper from Amsterdam, ‘How Do LLMs Cite?,’ traced which internal components decide to attach a citation, and found the decision leans on shallow heuristics, mostly the model noticing the same name shows up in the answer and in some candidate source.[3] Not ‘this document shaped my reasoning.’ Just: these tokens rhyme, good enough.[5]

And a Stanford benchmark called ClashEval found something that kills the obvious fix. Feed a model retrieved text that contradicts what it already knew, and it’ll throw out its own correct answer and adopt the wrong retrieved one more than sixty percent of the time.[4] So retrieval doesn’t reliably correct the model, and the model doesn’t reliably catch bad retrieval. The two problems don’t cancel. They stack.

If you want one number, the Tow Center at Columbia ran sixteen hundred queries across eight AI search tools, asking each to name the source of a quote it was handed. Easiest possible task. They were wrong more than sixty percent of the time, and they almost never admitted it, they just gave a wrong answer with a straight face.[6] The paid tiers were more confidently wrong than the free ones. Some invented links to pages that 404’d. And the detail every publisher signing a deal right now should tattoo somewhere: a licensing agreement bought no protection against being cited wrong. You can sell them your whole archive and still get misquoted in the answer about you.

It isn’t one web

Now the original hunch. The systems aren’t reading the same internet, and the reasons are more deliberate than I’d assumed.

First thing to get straight: ‘the AI crawled my site’ isn’t one event. Each company runs several bots doing different jobs. There’s the training crawler, gathering text to teach the next model. And there’s the retrieval crawler, grabbing pages in real time to build the answer in front of you. OpenAI runs GPTBot to train and OAI-SearchBot to answer. Anthropic runs ClaudeBot to train and Claude-SearchBot to retrieve. Separate doors. Separate locks.

Almost nobody outside the field has absorbed what this means. Block GPTBot because you don’t want to feed the training machine, and you’ve done nothing about whether you show up in ChatGPT’s live answers, which come through a different bot entirely.[7] A site can think it opted out of AI and still be sitting right there in the answer layer. Or think it’s fine and have quietly locked itself out of the only surface that still sends it readers. Everyone argues about the training door. The door that decides whether you appear in answers sits off to the side, unwatched.

Then there’s the asymmetry that explains my Gemini hunch. The Reuters Institute at Oxford measured which crawlers news sites block. By late 2023, roughly forty-eight percent of top news sites blocked OpenAI’s crawler. Only about twenty-four percent blocked Google’s. In the US it was worse, near seventy-nine percent blocking OpenAI against forty percent for Google.[8] Nearly every site that blocked Google also blocked OpenAI, but not the other way around. Google gets into rooms OpenAI can’t.

Figure 2. Publishers fence out OpenAI far more than Google, because blocking Google’s AI means losing Google Search too.

Why wave one through and slam the door on the other? Because Google fused its crawlers. The bot that indexes you for normal Google Search is tangled up with what feeds AI Overviews. Cloudflare’s CEO put it flatly: you can’t opt out of one without opting out of both.[9] Block Google’s AI and you vanish from Google Search, and no publisher living on search traffic can eat that. So the door stays open. Google’s edge isn’t only better engineering. It’s that it took the web’s biggest source of referral traffic hostage and made AI access the ransom.

Add what Gemini does at answer time. It grounds responses in Google’s live index, the actual one, which no competitor can copy because no competitor has it. A model with no live grounding is frozen at its training cutoff and will hand you last year’s Emmy winner without blinking. A model grounded in someone else’s index does better. A model grounded in Google’s own, while getting blocked half as often as its rivals, does better still. That’s the whole hunch, and it survives contact with the evidence: Gemini’s apparent edge on current events comes from owning the index and being let through more doors.

I won’t let myself off easy on this, though. That same Tow Center study found Gemini among the worst at citing accurately and most likely to fabricate a link. Fresh and well-sourced are different axes. Gemini is often the first and frequently not the second.

The fencing is getting more aggressive too. In July 2025 Cloudflare, which sits in front of a big chunk of the web, switched new domains to block AI crawlers by default. Opt-in instead of opt-out. They bolted on a pay-per-crawl system that revives the old ‘Payment Required’ status code so sites can charge bots.[10] Around the same time they accused Perplexity of using stealth crawlers with rotating addresses and faked identities to reach content after its declared bots were blocked, even on fresh test domains nobody should have known existed.[11] Perplexity called it user-driven activity, not crawling. Whatever you think of that particular fight, the trend is obvious. The freely crawlable web is being fenced off parcel by parcel, unevenly, in favor of whoever already had the leverage.

So three assistants, one question, three different subsets of the web, each carved out by who blocked whom and who signed what. The differences in their answers aren’t noise. They’re a map of a commons coming apart.

The answer is partly about you

Here’s a layer I hadn’t even thought to be suspicious of. The answer isn’t only shaped by which web the system can reach. It’s shaped by what the system thinks it knows about you.

Memory turned on quietly across most of the assistants. A 2026 study looked at real ChatGPT users’ stored memories and found about ninety-six percent of them were created by the system on its own, not by anyone choosing to save something.[12] People were being modeled constantly, mostly without opting in, and a real chunk of those memories held the kind of personal data privacy law treats as sensitive. The product copy says you’re in control. The numbers say the control is mostly set dressing.

Why it matters for what you find out is the obvious thing, and the researchers were careful to note the precise effect on news hasn’t been isolated yet. But the shape is plain. If the answer gets tuned to a model of who you are, two people asking the identical question get slightly different worlds. The filter bubble shows up in the answer engine, except now it’s wearing the costume of neutrality. People trust the AI because they think it’s objective, right as it’s personalizing underneath them.

What counts as trustworthy now

If the old web had a trust signal it was some messy pileup of links and reputation, the slow business of getting cited by other people. AI answers are renegotiating all of it live, and the early data is strange.

Pew looked at tens of thousands of real searches and found the sources cited most in Google’s AI summaries were Wikipedia, YouTube, and Reddit.[13] Forum posts and videos, in other words, are now a big share of what the machine treats as worth quoting. And it splits hard by engine. Reddit might be a quarter of Perplexity’s citations and a rounding error on Gemini.[14] Same question, different engine, different idea of where truth lives. One thinks the answer’s in a forum thread. One thinks it’s on a government page. Both say it with the same confidence.

Figure 3. The same query routed through different engines pulls from different worlds of sources.

This is the ground my own work stands on. Generative engine optimization, answer engine optimization, whatever name lasts. It’s real enough that enterprises were reportedly aiming around twelve percent of their digital marketing budgets at it in 2025, nearly all of them planning to spend more.[15] Old-school SEO still moves Google’s surfaces because Google reuses its ranking systems, but the other engines reward different things, clean extractable facts, answers up top, a presence in the community sources they happen to favor. We’re all learning to write for the thing that summarizes us instead of the person who used to click.

And where good information is thin, something worse happens. Researchers call them data voids, topics where solid sources are scarce, so the model grabs whatever’s there, a content farm or a propaganda site. A Harvard study found this is mostly mundane, a supply problem more than a successful manipulation campaign, which is reassuring until you notice the implication.[16] On exactly the questions where you most need a reliable answer, the new ones, the obscure ones, the ones nobody good has covered yet, the system is at its most gullible.

The bottleneck, and who’s watching it

Pull back far enough and one shape comes out of all of this. A few companies are setting themselves up as the layer the world’s information passes through, and the economics under that layer are hollowing out the thing it draws from.

Pew found that when an AI Overview shows up, people click through to a real source far less, and a quarter of those searches end the session right there.[17] Nobody leaves. The answer was enough. Other measures put the share of searches ending without a single click somewhere near seventy percent.[18] The traffic that paid for the open web is draining, and the trade is lopsided to a degree that’s hard to believe, some AI systems scraping thousands or tens of thousands of pages for every visitor they send back.[19]

Figure 4. When the answer sits at the top of the page, the click that funded the open web mostly doesn’t happen.

Figure 5. The old search bargain sent readers back. The new one mostly keeps them.

The old search bargain, let us index you and we’ll send readers, got replaced by one that indexes you and keeps the reader.

This isn’t a fringe complaint anymore. A US federal judge ruled in 2024 that Google is a monopolist in search, and the remedies fight treated AI as central to where the market goes next, though it pointedly didn’t hand publishers any more control over how AI Overviews use their work.[20] Brookings argues the case shows competition policy needs rebuilding for this era, and separately that the AI licensing market is busy rebuilding the exact gatekeeping it was supposed to break, with deals only the biggest publishers can reach and a citation premium that’s already gone.[21] Mozilla has funded work warning that compute and data are concentrating to where only the incumbents will be able to build the best models, because only they can afford the data.[22] The News/Media Alliance formally asked US regulators to step in before, in their phrase, the damage becomes irreversible.[23] A European publisher coalition filed an antitrust complaint over AI Overviews, saying they have no real way to opt out without disappearing from search, pointing to traffic losses they put as high as ninety percent.[24]

So, my opening questions, answered straight. Yes, people are doing serious peer-reviewed research, and it says the black box is real and the citations don’t hold. Yes, organizations are worried, and the worried ones keep turning out to be the credible ones, the universities and regulators and the publishers watching their own numbers fall. Yes, there’s a bottleneck forming, and it’s not a figure of speech, it’s a measurable shift in who controls what we can find out. And no, most people haven’t clocked it, because the whole thing is built to feel like an upgrade. Faster, cleaner, friendlier. The cost is one layer down, in the sources nobody visits now, the publishers nobody pays, the pages dropped from the answer that you’ll never know were dropped, and the citation you trusted because it looked exactly like one.

I’m not a doomer here, and I’m not going to pretend the old web was some lost paradise, it had plenty wrong with it. But seeing clearly comes first. So, for anyone whose work or readership now runs through these systems:

Know the doors are separate.

If AI search matters to your visibility, learn the difference between the bot that trains and the bot that retrieves, and check that you didn’t lock the wrong one. The expensive mistake is shutting yourself out of the answer layer while thinking you only opted out of training.

Treat the engines like different countries.

What earns a citation in Google’s AI isn’t what earns one in Perplexity or ChatGPT. And since the model often cites by shallow matching, being the cleanest, most quotable source on the page is its own kind of authority.

Quit measuring only traffic.

The click is dying as a unit of worth whether anyone likes it. Share of voice in the answer, branded search lift, audience you actually own. The publishers who make it through this are the ones who stopped renting their readers from a middleman that decided to keep them.

Watch the policy fights.

They’re the only force big enough to reopen any of this. The appeals, the European complaints, the licensing battles, the pay-per-crawl experiments. Any one of them could move the fence.

SOURCES

7. OpenAI documentation: sites opted out of OAI-SearchBot are not shown in ChatGPT search answers. Robots.txt itself is advisory (RFC 9309).

13. Pew Research Center analysis of 68,879 Google searches by ~900 U.S. adults, 2025. Wikipedia, YouTube and Reddit were the most-cited sources; .gov links were 6% of AI-summary links.

14. Tinuiti ‘AI Citations Trends Report Q1 2026’ (with Profound); Reddit share figures by engine. Engine-level citation behaviour varies widely and shifts quarter to quarter.

15. Conductor, ‘The State of AEO/GEO in 2026’ (via eMarketer): U.S. enterprises averaged ~12% of digital-marketing budgets on GEO in 2025, with ~94% planning to increase spend.

16. Alyukov et al., Harvard Kennedy School Misinformation Review, Oct 2025 (the ‘data voids’ finding); cf. NewsGuard’s Mar 2025 audit reporting a higher one-third repetition rate.

17. Pew Research Center, Jul 2025: with an AI Overview present, users clicked a traditional result 8% of the time vs 15% without; 1% clicked a link inside the Overview; 26% ended the session. Google has disputed the methodology.

18. Similarweb: zero-click searches rose from ~56% to ~69% between May 2024 and May 2025.

19. Crawl-to-referral ratios: TollBit, Q1 2025 (OpenAI 179:1, Perplexity 369:1, Anthropic 8,692:1). Google Search baseline (~14:1): Cloudflare, 2025. Different sources and windows; treat as orders of magnitude, not precise.

20. United States v. Google LLC, Judge Amit Mehta (D.D.C.), liability ruling Aug 2024; remedies decision Sep 2025. Google has said it will appeal.

NEXT PROJECT

DISCOVER MONEY TOOLS

blurry female portrait

LET'S BUILD
SOMETHING GREAT

Whether you have a project in mind, a question, or just want to talk AI and UX, I would love to hear from you.

blurry female portrait

LET'S BUILD
SOMETHING GREAT

Whether you have a project in mind, a question, or just want to talk AI and UX, I would love to hear from you.

blurry female portrait

LET'S BUILD
SOMETHING GREAT

Whether you have a project in mind, a question, or just want to talk AI and UX, I would love to hear from you.