Popular Posts

Latest Articles

Semrush Partner

How GEO Actually Works: 6 Myths to Drop

This week I’m not reporting news, I’m writing a guide. A language model decides what to say about your brand based on only two things: training data and grounding. Almost…
how geo actually works - 6 myths to drop

This week I’m not reporting news, I’m writing a guide. A language model decides what to say about your brand based on only two things: training data and grounding.

Almost every myth in our industry comes from missing that distinction. I take the six biggest ones apart. At the center sits the llms.txt saga, and its real lesson has nothing to do with the file: don’t take Google’s word for it, and don’t take your favourite expert’s word for it either.

Then one live observation: I asked Google’s AI which credit card is best. It named seven banks by product. Not one of the sources it cited was a bank.

Hi everyone!

A reader asked me a question last week that stopped me: “Fine, we measure AI visibility. But what is the model actually looking at when it talks about me?”

It’s a simple question, and it sits at the exact center of the confusion I see in the field. Most GEO conversations are built on a wrong picture of how a model produces an answer. And a wrong picture never produces a right strategy.

So this week: no news roundup. Just the mechanism, the six myths that fall out of getting it wrong, and one live test from a market most of you never look at.

Grab your coffee, let’s dig in.

The mechanism: a model knows you in only two ways

Bottom line up front (BLUF): Only two mechanisms decide what a language model says about your brand, and nearly all of your leverage sits in the second one.

1. Training data. A frozen snapshot of the internet, taken the moment the model was trained. If your brand was in that snapshot, the model knows you. If it wasn’t, it doesn’t. The bad news: you can’t easily change this. Influencing future training cycles takes months.

2. Grounding. The moment a user asks a question, the model goes out, fetches live sources, reads them, and writes the answer from what it found. The technical name is RAG (retrieval augmented generation); the plain name is web search. Engines typically scan 40-50 sources for a query and pull the 12-20 most relevant into the answer.

Here’s the good news: grounding is dynamic, and it is entirely your playing field. A model can cite a piece published this morning. It can recommend a brand it has never seen in training, as long as that brand is well represented at retrieval time.

So “we’re small, AI doesn’t know us” is not a fate. It’s a starting position.

The 6 myths the industry still believes

Myth 1: “If the model doesn’t know me from training data, it’s over.”

It isn’t. Training data is a frozen photo; grounding is a live feed. Even if your name never appeared in training, you enter the answer if you’re findable at the moment the question is asked.

What to do: Stop spending energy trying to “teach the model who you are.” Spend it being present where it looks when the question is asked.

Myth 2: “Models use Google, so ranking in Google is enough.”

The most expensive myth on this list. Google’s AI Overviews does use Google, true. But ChatGPT is partly fed by Bing, Perplexity runs its own index, Claude uses Brave, and Grok scans X alongside web search.

What to do: You don’t need to optimize separately for all of them. But make sure you haven’t accidentally blocked any of them. Being number one in Google while remaining completely invisible to ChatGPT is entirely possible today. Open your robots.txt this week and look.

Myth 3: “I’ll optimize for one query.” (And its twin: “I’ll build a page for every fan-out.”)

The model does not run your question as written. It splits a single prompt into 2, 6, sometimes 20 separate searches. This is called fan-out. “Best CRM for a 50-person B2B company” fans out into “best B2B CRM 2026,” “CRM for mid-market SaaS,” “CRM pricing under 50 users,” and so on.

Teams that learn this then fall into a second trap, and Lily Ray wrote about exactly that this week: chasing individual fan-out queries is a poor use of time, and building one page per fan-out will likely land you in SEO spam territory. In her words, it echoes the pre-Helpful Content Update playbook. The right move is to analyse fan-outs at scale and cluster them into core topics.

She adds a caveat I agree with: keep SEO tools with real search volumes in the process, rather than invented “prompt volumes.”

What to do: Don’t build for a single fan-out. Build for the topic cluster behind the prompt. And watch the words the model adds during fan-out; the two most common additions are the current year and “reviews.”

Myth 4: “If my page was cited, my page was read.”

No. The model doesn’t take your whole page, it takes a passage. When a 3,000-word article gets cited, the thing that actually entered the answer may be a single 40-word paragraph.

What to do: Put a self-contained, standalone answer paragraph near the top of every important page. Bury the answer in paragraph 14, between three conditional clauses, and the model can’t extract it.

Myth 5: “I’ll drop in a magic file and be done.” (This is the heart of the guide.)

Google, in its guide to optimizing for generative AI features, under a section literally titled “mythbusting,” wrote that you don’t need machine-readable files like llms.txt to show up in AI search. The industry said “case closed.” I’ve been saying “stop chasing magic files” for months myself.

Then Google contradicted itself. Chrome’s Lighthouse made its Agentic Browsing audit a default in May 2026, and that audit checks whether your site has an llms.txt. One hand says “don’t bother,” the other hand goes looking for the file.

Then Cyrus Shepard (founder of Zyppy, former Head of Global SEO at Moz) stopped talking and ran a test. What he found: Google recognizes llms.txt files, reads what’s inside, and can cite them as a source in its answers. In one of his examples, the only URL cited for that topic was the llms.txt file itself. So “I completely ignore it” is not, strictly speaking, true.

Now flip the coin. Experts raised a fair objection: this may not be special treatment for llms.txt at all. Google already crawls and indexes any internally linked .txt file; whether it’s named llms.txt or notes.txt makes no difference to a crawler. Cyrus explicitly agreed with that technical explanation.

So the correct reading is: Google doesn’t ignore llms.txt, but it doesn’t bless it as a standard either. The truth sits somewhere between the two extremes.

Meanwhile, Ahrefs studied 137,000 sites and settled the practical question: 28% of sites publish an llms.txt, but 97% of those files received zero requests in May 2026. AI retrieval bots (the ones actually generating live answers) accounted for just 1.1% of all requests. And no AI bot ever went looking for an llms.txt that didn’t exist. Nobody is knocking on that door.

The real lesson here goes far beyond llms.txt: when Google says something, don’t take the plunge. “I ignore it” turned out not to mean “it’s entirely invisible.”

But don’t reserve your skepticism for Google alone. Cyrus himself was once among those saying “there’s little evidence for llms.txt,” then he tested and updated his position. That’s a good thing; a good expert changes their mind with data. But the lesson for us is this: the people we admire can be wrong, and they can reverse. Don’t accept Google’s statement or an expert’s current sentence without question. Your own test is the most reliable compass you have.

One balancing note: none of this means llms.txt is a Google ranking factor. It isn’t.

Myth 6: “If I correct the error once, the model will learn.”

It won’t. The model isn’t learning. It grounds again on every query, going back to the same pool and refreshing the same information.

Here’s my own data. Across a five-week tracking study, the error giving Stradiji’s founding year as “2010” (it’s 2009) repeated in four of five engines, in all five weeks. Even though the correct information sits plainly on our own pages. The reason is simple: as long as the third-party pages feeding that error remain, the model rediscovers the same mistake every single week.

What to do: Correcting your own page isn’t enough. Correct the sources feeding the error.

The next test: OKF (Open Knowledge Format)

I want to raise this right after myth #5, because the next test of our discipline is already here.

What is OKF? Open Knowledge Format, announced by Google Cloud on June 15, 2026, is a proposed open standard for packaging organizational knowledge in a format AI agents can read. Technically it’s plain: markdown files with YAML frontmatter. Each concept (a dataset, a metric definition, an API endpoint) becomes a document, which Google calls a “tree”; bundle the documents together and you get a “forest.” Vendor-neutral, version-controllable, readable by both humans and AI. The spec is open on GitHub.

Google’s rationale is striking: as foundation models keep improving, the thing limiting them is no longer intelligence, it’s lack of context.

Why it matters. The center of gravity in AI competition is shifting. The game is no longer about building a better model; it’s about feeding the model the right context. Which is really just the new version of something SEOs have known for years. Having the best content was never enough on its own; the knowledge had to be structured, connected, and made machine-readable. That’s why we’ve been investing in structured data and entity SEO all along. OKF is the same problem in the agent era.

Now the discipline, pay attention. This is where the whole lesson of this issue applies: OKF is not a Google ranking factor. Dropping an okf/ folder on your site will not make you more visible in AI Overviews. OKF was designed for internal knowledge layers and agents, not for the open web. And Google itself writes plainly: “OKF v0.1 is a starting point, not a finished standard.”

So let’s not watch the llms.txt movie again. The industry’s reflex is predictable: “Google shipped a new format, let’s generate one, maybe it helps AI visibility.” It won’t. At least not today, not on the evidence we have.

What to do instead? Invest in the principle underneath the format, not the format itself: knowledge a machine can’t read is knowledge that doesn’t exist for an agent. That principle holds whether OKF wins or dies. Take your organizational knowledge (product data, pricing logic, your areas of expertise, your FAQs) out of scattered PDFs and decks nobody opens, and move it into a structured, current, readable layer. Do that and you’re ready for whichever standard wins. Skip it and no file will save you.

One live observation (an observation, not a study)

Here’s where I’d normally show you a ranking table from an emerging market. I’m not going to, and the reason is itself the point of this issue.

I pulled ranking data from SEO tools for this section. Striking tables came out. I nearly led with “this site ranks #1.” Then I ran the same queries live in my own browser, and the results didn’t hold. Rankings move with session, personalization, location, and exactly how you phrase the query. What the tool shows and what you see are not the same thing.

So no ranking claims from me. Instead I’m looking at something far more solid, and far more important: citations. That’s the visible face of grounding.

I asked Google: “which credit card is best?” (in Turkish, in Turkey). AI Overviews gave a competent answer, naming seven banks by product: Garanti BBVA Shop&Fly and Miles&Smiles, İş Bankası Maximiles, Yapı Kredi Worldcard, Akbank Axess, and QNB, TEB and DenizBank for cash advances.

Then I opened the sources: Cimri, HesapKurdu, getkampania, HangiKredi, TeklifimGelsin. All comparison and deals sites. I asked Perplexity the same question; its visible citations were a loan blog, a YouTube video, and again a comparison site.

Two independent engines, same picture: the answer describing the banks’ own products is not being fed by the banks.

It gets worse. Google’s AI Overview closes by pointing the user to HangiKredi and HesapKurdu to compare cards and apply directly. Google’s AI is sending the conversion not to the bank, but to the intermediary.

And the most instructive detail. Exactly one bank page made it into the citations: Yapı Kredi. Why that one? Because that page doesn’t say “we’re the best.” It documents something: “the first and only Turkish credit card selected for The Banker’s list of the world’s 10 most valuable credit cards.”

The page that praises itself doesn’t get cited. The page that documents someone else’s praise does.

Proof beats posture. We’ve been saying it for years; here it is, tangible.

Now the limits, because that’s the whole theme this week: this is one query, one day, one session. I’m not concluding anything about a national market. What I want from you isn’t to read my table. It’s to run the same test in your own category.

Checklist (do this week)

  1. Look at the sources, not the answer. Ask an engine your target question, then open the source list underneath. Are you in it, or only people writing about you?
  2. Run the same test on at least two engines. If the citations overlap, you have signal. If they don’t, you’re still looking at noise.
  3. Never trust a single observation. Repeat on different days, in different sessions, logged out. A finding that doesn’t repeat isn’t a finding.
  4. Check the doors. Is GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Google-Extended access open in your robots.txt? Are you indexed properly in Bing?
  5. Don’t chase fan-outs, cluster them. Group queries into core topics and compare them against your content portfolio.
  6. Publish proof, not posture. Instead of claiming “we’re the best,” document an independent ranking, award, or third-party review. That’s usually the page AI finds worth reading.
  7. Treat intermediaries as a channel. If the text the model reads sits on a comparison site, being represented accurately there isn’t an SEO job. It’s a distribution job.

The week in brief

  • Search is turning from answer into conversation. Google opened up follow-up questions directly from an AI Overview, flowing into AI Mode, and expanded personal intelligence to roughly 200 countries and 98 languages. Users don’t read an answer and leave anymore; they keep talking. Which makes optimizing for a single query even less meaningful.
  • Ads are moving inside the answer. ChatGPT Ads can now auto-generate ads, added audience targeting, and expanded to Japan and South Korea. Remember the credit card example: the sources under the answer were already intermediaries. Now an ad layer is entering the answer itself. The space for a brand’s own voice keeps narrowing.
  • A model traffic jam. OpenAI made GPT-5.6 generally available on July 9 (Sol, Terra, Luna) and it’s now ChatGPT’s default; Axios reports ChatGPT Work landed the same day. xAI took Grok 4.5 public on July 8. Google pushed Gemini 3.5 Pro to July 17. Why this matters to you: when the default model changes, the answer most users see changes. Reasoning depth changes, fan-out behaviour changes, and therefore which sources get pulled changes. Your June AI visibility measurement may already be void. Re-measure.
  • Cloudflare’s “content-signals” tag does nothing. John Mueller confirmed on July 6 that no crawler or LLM uses it. A fresh rerun of myth #5.

Closing

The essence of this guide in one line: AI doesn’t describe you from memory. It describes you from whatever sources land in front of it at that moment. Training data isn’t in your hands. Grounding is entirely your field.

But the thing that stayed with me this week isn’t technical. It’s this: nobody in this business hands us the truth pre-packaged. Not the platforms, not the tool vendors, not us consultants. Google said “I ignore it,” and it wasn’t quite so. An expert we all respect said one thing, tested, and changed his mind. A new standard just shipped, and the industry will predictably sprint at it.

And to be honest with you: I threw away this issue’s data three times. I had striking ranking tables in hand. I checked them live, they didn’t hold, and I deleted every one of them. Filling a newsletter is easy. Filling it correctly sometimes means binning what you’ve got.

So the real lesson this week isn’t technical, it’s professional: our job is not to believe what we’re told. It’s to question it and run our own test. Myself included. Test everything you read here against your own data.

Now my open question to you: have you ever asked an AI engine “who’s the best in my category” and then actually opened the source list underneath the answer? In banking, that list had almost no banks in it. What shows up in yours?

Share!

These May Also Interest You

Craving more SEO knowledge? Extend your learning with #SEOSDINERSCLUB

Subscribe to our newsletter for weekly SEO insights, join the discussion in our community, or engage with professionals on our Twitter group.