It penalizes commodity, and scale decides which one you made. Two axes to place your content on, plus what a 220-site study shows about scaling AI output.
The panic that started with Claude’s text watermark turned into “write with AI and you’re finished.” Google’s own documentation says the opposite in its very first sentence: generative AI is useful for researching a topic and for adding structure to original content.
What gets penalized isn’t the tool, it’s producing many pages without adding value.
A viewer on my YouTube channel asked a sharper question than the industry has been asking, so this week I answer it on two axes: do you have original data, and at what scale are you publishing it? Lily Ray’s 220-site study shows what scale does. Then I add my own: an AI engine wrote out its own reasoning for shortlisting us, and it never once asked who wrote the content.
I didn’t pick this week’s topic, you did. A comment from Yıldırım Sertbas on my YouTube channel frames the question better than most of the industry has managed:
“Can we get a video about AI content? What is Google actually penalizing? Content woven from data we produced ourselves, data AI could never generate, but where AI built the structure? Or AI content that just blends what’s already on the internet, the kind anyone could produce? And of course whether it’s been scaled to high volumes matters too.”
When I read it I thought: he has already split the answer into three. Original data, blended information, scale. While most of the industry is still arguing about whether using AI is allowed, this question grabs the right end of the stick. Thank you, Yıldırım. This issue belongs to him. Grab your coffee.
The thesis: it’s not the tool, it’s commodity and scale
Google doesn’t penalize AI content, and it never has. What it penalizes is content produced at scale without adding value for users.
I’m not inventing this distinction; Google’s own documentation states it in the opening sentence: “Generative AI can be particularly useful when researching a topic, and to add structure to original content.” So the first option in that question, content woven from our own data with AI building the structure, is the use Google itself recommends.
The sentence that follows describes the second option: “using generative AI tools or other similar tools to generate many pages without adding value for users may violate Google’s spam policy on scaled content abuse.”
Read it carefully. It does not say “pages made with AI.” It says “many pages without adding value.” Two variables: value and volume. The tool isn’t in the equation.
Here’s the practical frame I use: place your content on two axes. Do you have data nobody else has, and at what scale are you publishing it? The risk lives at the intersection.
First, let’s close the watermark question
Anthropic lit the fuse: Claude models released after 2 August 2026 embed an invisible statistical watermark in the text they generate. The purpose isn’t punishment, it’s regulation. The EU AI Act’s transparency rules took effect the same day.
Two mistakes are being made here. First, the watermark doesn’t tell you Claude wrote the text, it tells you the text passed through Claude. Edit your own writing with it and a trace can remain. So it’s a record of processing, not proof of authorship.
Second and more important: Google’s position hasn’t changed since February 2023, and it is blind to production method. A watermark is a detection tool, not a ranking signal. Confusing the two is like mistaking the thermometer for the fever.
Which writer today builds an outline without these tools? The real question is whether what comes out has something in it that nobody else has.
What Google’s own text says
Let’s skip everyone’s commentary and go to the source. Google’s generative AI guidance (developers.google.com, last updated 10 December 2025) says, briefly: AI is useful for research and for adding structure to original content; using it to produce many pages without adding value counts as scaled content abuse; focus on accuracy, quality and relevance when generating automatically; give readers context about how the content was created.
For the deeper version it points to sections 4.6.5 and 4.6.6 of the Quality Rater Guidelines.
Look at that 4.6.6 definition: main content created with little to no effort, little to no originality, and little to no added value. The words “artificial intelligence” appear in none of it. A human can absolutely meet all three criteria, and plenty do.
The scaled content abuse policy introduced in March 2024 contains the decisive phrase too: no matter how the content is created. Google’s target isn’t who wrote the page, it’s why the page exists.
A two-axis answer to the question
Put the three elements from that comment into a single frame and the answer falls out. Two axes: do you have data nobody else has, and at what scale are you publishing it?
Original data YES + low scale = Safe zone. Let AI build the structure and draft it, no problem. This is exactly the first option in the question.
Original data YES + high scale = Legitimate but hard. If every page genuinely carries first-party data, it’s legitimate. If the template is identical and only the data isn’t, it slides fast.
Original data NO + low scale = Invisible zone. You won’t be penalized, but you won’t rank either. The problem isn’t punishment, it’s irrelevance.
Original data NO + high scale = Red zone. This is the textbook definition of scaled content abuse. It’s the second option in the question, multiplied by scale.
The most overlooked of these four is the third: no original data, but low scale too. People read that situation as “I got penalized.” There is no penalty. That content simply doesn’t show up, because ten pages already rank for the same information and yours adds nothing new. A penalty is an enforcement action; invisibility is an outcome. Most brands aren’t penalized, they’re just producing commodity.
The second cell deserves an answer too: “I run an ecommerce site with thousands of product pages, what about me?” If every page carries real data specific to that product (price, stock, specifications, genuine user reviews), that’s legitimate scale. The trouble starts on pages where the template stays identical and only the brand name changes.
Why scale is decisive: the 220-site evidence
The most concrete evidence is Lily Ray’s May 2026 study. She tracked more than 220 sites that AI content platforms listed as customers in their own published case studies. Measured against peak organic traffic:
Loss against peak traffic
Share of sites
Lost 30% or more
54%
Lost 50% or more
39%
Lost 75% or more
22%
Source: Lily Ray, “It Works Until It Doesn’t”, 13 May 2026. 220+ domains, third-party estimates from Ahrefs and Sistrix. In her own words this is a correlation she observed, not a causal claim.
The recurring shape: rapid growth in page count over six to twelve months, then a traffic peak, then a steep decline within the following year that erases most of the gain. Glenn Gabe calls it “Mount AI”: steep up, steep down.
Most of the declines happened after the case studies were published, and today some of those case studies are still live while the pages they celebrate have been removed from the sites.
Lily Ray also lists eight templates she considers risky: comparison pages at scale, “what is X” glossaries, “best X for Y” listicles, self-promotional listicles where the publisher names itself number one, a dedicated “alternatives” page for every competitor, programmatic location and language scaling, FAQ farms with one question per URL, and off-topic content at scale.
The critical nuance is hers: these page types do work. The problem is that once scaled, they become a detectable footprint.
The item on that list I keep coming back to is FAQ farms, because for the past year this is exactly the structure people have been recommending “so AI will cite you.” A reminder: on 7 May 2026 Google quietly deprecated FAQ rich results, without even a blog post, just a note added to the documentation. So the structure built to win citations lost its support within months. That’s the cost of chasing tactics.
Where the experts converge: human oversight
Lily Ray‘s sentence is the one the industry needs to hear this year: “The tools themselves are not the problem, but the implementation can be.” My favourite of the questions she suggests asking before you publish: could a competitor publish a near-identical version of this page tomorrow using the same prompt? If yes, that page is commodity.
Aleyda Solis arrives at the same place from the technical side: leveraging AI is possible, but it requires a personalized editorial and optimization workflow to ensure quality, originality and expertise, integrating unique brand insights and first-party data. Her point lands hard: that is precisely what AI platforms are likely to cite.
Cyrus Shepard builds the GEO bridge: because AI answers are grounded in real search results, ranking in Google still decides whether you get cited. Ranking on page one for a single query gives you roughly a one-in-three shot at a citation; ranking across the fan-out queries pushes that to 80-90%.
Marie Haynes points to the wording change in Google’s guidance: “content written by people” became “content created for people.” That one-word edit ends the argument. The criterion isn’t the type of author, it’s who the content exists for.
The common denominator is a single word: oversight. AI builds the draft, a human takes responsibility.
My own data: the engine wrote out its own reasoning
For twelve weeks I’ve been tracking my own brand across five AI engines. In this week’s audit, when I asked ChatGPT about enterprise GEO consultants, it thought for 41 seconds, scanned 22 sites, and crucially wrote out the basis for its ranking itself: concrete published GEO case studies, enterprise SEO history, technical capability, current AI-search research, and visibility in third-party sources.
Its stated reason for shortlisting us wasn’t our services page, it was the research I published throughout 2026.
Don’t miss the lesson. Not once did the engine ask “who wrote this content, did they use AI?” The only thing it asked was what we had proven.
And in the same audit my position slipped from 2 to 3, not because we weakened: a competitor entered the pool with a published, numbers-backed case study. What beat me wasn’t more content, it was more measurable proof.
Is the industry building that proof layer? No
This week I went back to my 60-site audit from early August and looked at it through the “proof and human layer” lens:
Missing signal
Sites
Share
No author information or E-E-A-T signal
25/60
42%
No independent review or reputation source
31/60
52%
No Wikidata entry
52/60
87%
Source: my own GEO Score Card archive, 60 sites, 28 April – 7 August 2026. Aggregate and anonymized.
So on nearly half of these sites there is no signal at all about who wrote the content or who stands behind it. Forget the argument about whether AI wrote it or a human did: nobody has signed it.
The emerging-market lens
I think this risk runs higher in markets like mine. In Turkish, the cost of scaling content with AI has fallen to nearly zero, while the culture of producing original data (publishing your own measurements, running an industry survey, writing up a first-party case) is still very thin. Put those two together and you get a highway straight into the red zone.
An honest note: of those 60 sites, roughly 46 are from Turkey and 14 from abroad, and the same gaps showed up in the international ones too. So a thin proof layer is a maturity problem, not a national one. What distinguishes emerging markets is simply that scaling costs even less here.
The upside: when competition is this weak, the first brand to publish its own data stands alone in its category.
A caveat: what I am not saying
Lily Ray’s data consists of third-party estimates, and she explicitly writes that it’s a correlation, not a causal claim. My 60-site data is an aggregate audit observation, not a controlled experiment; my ChatGPT observation is a single personalized scan.
So I’m not saying “54% of everyone using AI for content loses traffic,” and neither should you. What I am saying: Google’s policy targets value and volume, not method, and field data keeps showing scaled commodity arriving at the same destination.
What to do this week: three questions before you publish
Don’t produce anything new this week. Take what you published in the last month and ask it these three questions.
Originality: Is there a single thing on this page that isn’t already in the top ten results? Your own measurement, your client outcome, a test you ran. If not, you’re producing copies, not content. The shortcut: could a competitor publish the same page tomorrow with the same prompt? If yes, don’t publish it, fill it with data.
Scale: How many of this format exist on the site? A handful, or hundreds? Hundreds of copies of the same template turn individually harmless pages into a collective footprint. If you’re running three or more of those eight templates together, stop and look.
Accountability: Whose name is at the bottom of this page, and did that person actually read it? On 42% of the sites I audited there wasn’t even an author signal. It’s the cheapest gap to close.
One sentence: use AI for the draft and the structure, supply the proof yourself, and let a human sign it. Get those three right and nobody cares what tool you used.
Join the benchmark
If you’d like me to map where your content sits on these two axes and where the gaps are in your proof signals, that’s exactly what the GEO Score Card is for: I produce a free AI visibility roadmap covering five engines for every brand that applies.
Mert Erkal is the founder of Stradiji, which has been providing consultancy services on Search Engine Optimization (SEO), SEO Friendly Content Production and Optimization, and Conversion Optimization since 2009. SEO consultancy of enterprise companies is Mert's unique expertise. He has been sharing and commenting on weekly critical developments from the SEO world for about three years with his newsletter "SEOs Diners Club." With the advantage of remote working, he continues to provide SEO consultancy to English-speaking countries, especially the United States, Australia, and the United Kingdom.
LinkedIn'de Takip Edin