“ChatGPT uses Google or Bing” got corrected this week, and then argued with. The Peec AI team documented ChatGPT’s own family of indexes: web, PDF, YouTube, news, Wikipedia, local, finance, legal, medical, shopping and more.
The evidence runs from sworn court testimony to job postings, from live A/B tests to a billion-page crawl experiment.
But the objections underneath the post are serious too, and almost nobody read them. I read both.
I picked this week’s topic from a LinkedIn thread and the article behind it. Here’s why: the post itself was interesting, but the objections underneath it were more interesting, and almost nobody read them.
Three weeks ago I shared this with you: when ChatGPT recommends brands, it doesn’t gather the candidates from the web, it writes them from memory. The most common question I got back was: “Fine, but when it isn’t working from memory, where does it actually look?”
This week that answer got a lot clearer. Grab a coffee.
“ChatGPT uses Bing or Google” is now an incomplete sentence
ChatGPT has built its own search index, and it isn’t one index but a family split by subject: web, PDF, YouTube, news, Wikipedia, local, finance, legal, medical, shopping and images.
Alongside it, ChatGPT still uses Google and Microsoft, drawing on at least eight separate providers.
The practical consequence: “I’m visible in Bing and Google, so I’m visible in ChatGPT” no longer holds. This is a well-evidenced reading, but I’m giving you both the evidence and the objections below, because the system changes week to week.
The librarian with twelve ledgers
Let me put this without the jargon. Think of ChatGPT as a librarian. You ask a question, it goes and finds sources.
For two years the assumption was that this librarian looked at somebody else’s catalogue: Bing’s. That’s where “ChatGPT uses Bing” comes from.
The picture today is different. Bing’s and Google’s catalogues are still on the desk. But there are also ledgers the librarian keeps itself. And it isn’t one ledger.
The list Peec AI documented: general web, PDF, YouTube, news (three separate tiers for the last day, the last week, and older), arXiv, Wikipedia, local, finance, legal, medical, shopping and images.
This is the same architecture Google spent twenty years building: one general index, with subject-specific verticals layered on top. So “being visible in ChatGPT” isn’t one job. If you sell products, the shopping ledger matters. If you’re a local business, the local ledger. If you publish health content, the medical one. They don’t behave the same way.
The evidence: four pieces, plainly
1. Sworn testimony. Nick Turley, who runs ChatGPT, testified under oath in the Google antitrust case. Search data from non-Google providers had “significant quality issues.” OpenAI wanted to work with Google at the time, and Google refused. Plan B was to build their own index.
Two dates in that testimony are instructive. Turley said they started in 2023, aiming to answer 80% of queries from their own index by the end of that year. But in the same testimony he added: even with full access to Google’s data, working out whether 100% is achievable at all would take at least five years.
2. Job postings. OpenAI’s listings ask for people who can “design and operate indexing systems, retrieval pipelines and serving layers.” One posting describes the team working at exabyte scale. An exabyte is a 1 followed by eighteen zeros. You don’t build for that to pass queries through someone else’s index.
3. A controlled experiment. This is the most convincing piece for me, because it’s reproducible. Malte Landwehr picked a site with zero organic traffic. On specific days he asked only ChatGPT about that site. On those days, the site’s Google Search Console showed a traffic spike. So ChatGPT really is querying Google while it builds an answer.
4. Live A/B tests. In mid-August an experiment named “prefer-index-over-serp-v3” was affecting 8% of chats. OpenAI is testing its own index against scraped search results on real traffic. As of 2 September there were at least five experiments running on the shopping side alone.
What the four say together: this isn’t a side project. It’s a deliberate build that has been running since 2023. And it isn’t finished.
And there’s a cache
This part may be the most immediately useful.
ChatGPT doesn’t fetch your page live on every question. It can’t: most web pages take longer than 0.8 seconds to load, and that isn’t workable at this scale. Instead it crawls pages in advance and stores them.
There’s an experiment that shows how aggressively. A researcher on the Peec team published a one-billion-page test site. As of early September, ChatGPT had crawled 6 million of those pages and was still going, at 35,000 requests per hour.
And you can test this yourself. ChatGPT’s settings include lockdown mode. In that mode it can’t open live pages; it can only show you what’s in its cache. So the question “is my most important page in ChatGPT’s memory?” can be answered in five minutes, without buying a single tool. That’s the most concrete gift in this week’s issue.
But the industry doesn’t agree
This section is worth more than the article itself, and I suspect you haven’t seen it, because the post travelled and the comments didn’t. The objections are serious and some of them didn’t get a convincing answer.
Chris Wheeler: “This is evidence about the plumbing, not about marketing strategy. ‘Figure out that source’ isn’t an instruction a team can test.” Nothing here establishes which intervention influences inclusion, how stable any effect is, or whether the resulting visibility reaches a commercially relevant buyer. Landwehr conceded: “The goal of the article was not to establish which interventions influence inclusion.”
Ryan Edwards: “Don’t think of this as the index itself. Think of it as a retrieval layer inside a broader index.” Landwehr accepted the correction.
Christoph Burseg: In one academic study, ChatGPT’s answers traced back only to OpenAI’s own crawler; neither Google nor Bing results appeared. Landwehr’s reply: “That study looked at something else. In recent weeks I’ve only seen Bing used for Deep Research.”
Juliane Bettinga: “What’s your evidence for a separate YouTube index?” The answer in the thread was an assumption. The article does list YouTube as a finding. Those are two different things, so go to the article rather than the thread.
Vincent Quero: “What are the weights, the proportions between these sources?” That one went unanswered.
Wheeler’s objection is the strongest, I think. The article partly answers it, since it does offer four concrete starting points. But “which intervention works” is still open.
And I’m applying one of his lines to this newsletter too: citation visibility is a diagnostic, not a success metric. Being mentioned more often by an AI does not mean more customers. Measure it, but don’t declare victory before you’ve tied it to revenue.
What’s certain, what isn’t
Certain: OpenAI has its own search infrastructure and it’s been a deliberate investment since 2023. Sworn testimony, job postings and live A/B tests point the same way together.
Highly likely: ChatGPT isn’t tied to a single engine; it switches source depending on the situation. The controlled Search Console experiment shows Google, the Deep Research logs show Bing.
Not yet proven: what proportion comes from which source, and which of our interventions gets us into these indexes. The second one is the part that makes money, and nobody has a clean answer today.
My own read: betting on a single engine is risky. “Let’s invest in Bing” is risky, and so is “Google is enough.” On top of that, the thing you assume is Bing may be arriving through Web IQ, a separate Microsoft product built for a different purpose.
What the model actually sees of you is still tiny
We don’t know who wins the engine argument. But one thing isn’t changing: ChatGPT is not reading your page end to end.
What it usually holds is your title, your URL and roughly the first 200 characters of your page. And that summary is not taken from your meta description. It’s taken from the start of your page body.
Rough arithmetic: your title + your H1 + about 150 characters left for you. The sneakiest thief is the alt text of the first image sitting under the heading; on its own it eats about 50 characters.
Worse: one in seven of the pages examined has no H1 at all. In that case your template decides where the summary starts, not you.
The emerging-market lens
Here’s the part of this story nobody discussed, and it matters well beyond Turkey.
The article says ChatGPT treats Yelp and TripAdvisor as licensed partners in its local vertical. Google Maps data reaches it too, but you have no say in how that gets sourced.
Now think about any market where Yelp barely exists and TripAdvisor only covers tourism and restaurants. Turkey is one; so are most of Latin America, the Middle East and South and Southeast Asia. In all of them, most local businesses are either absent from ChatGPT’s local ledger or represented there without any control over it.
There’s a door problem on top of that. Across the 60 sites I ran through our GEO Score Card between April and August, 44 (73%) had no explicit permission for AI bots in robots.txt, and 7 (12%) blocked them outright.
Whichever engine wins, the door is already shut on most sites. My priority order: bot access, then being in the cache, then what your first lines say, and if you’re a local business, your TripAdvisor listing.
The one step to take this week
Open your five most commercially important pages. Half an hour, no technical skill required.
- Turn on lockdown mode and ask for the latest version of your most important page. Are you in ChatGPT’s cache, and how old is the copy? This is the most concrete test of the week.
- Check OAI-SearchBot and GPTBot separately in your robots.txt. They do different jobs: one is search, the other is training data. Most sites conflate them and block both.
- Does the page have an H1? One in seven doesn’t. If it doesn’t, start there.
- Clear the runway above the H1. Dates, category kickers, breadcrumbs and especially that first image’s alt text are all eating your budget.
- If you’re a local business, check your TripAdvisor listing. It’s a licensed partner in ChatGPT’s local ledger, and in most emerging markets it’s the biggest gap.