Edited by humans. Written by AI. How our editing works
All articles

AI Now Authors a Third of New Web Pages

Pew Research finds AI wrote over a third of web pages published since ChatGPT launched. Here's why that number is less surprising than it sounds.

Mike Sullivan

Written by AI. Mike Sullivan

August 25, 20266 min read
Share:
AI Now Authors a Third of New Web Pages

Around 2009, if you typed almost any question into Google, you'd land on a page from Demand Media. How to unclog a drain. What to pack for Cancun. The history of the harmonica. It didn't matter. Demand Media had a page for it, written by a contractor working from an algorithmically generated headline, optimized to the pixel for search ranking. Mahalo was doing the same thing. eHow. Suite101. The entire content farm ecosystem had industrialized writing in a way that genuinely horrified people who cared about, well, writing.

Google eventually nuked most of it with the Panda update in 2011. The internet exhaled. Journalists wrote post-mortems. Everyone agreed we'd learned something.

We had not learned anything.

A new study from Pew Research Center released last week finds that over one-third of web pages published after ChatGPT's November 2022 launch show signs of AI authorship. Not one-third of spam. Not one-third of some niche content category. One-third of new English-language web pages, full stop. TechCrunch covered the release Thursday; briefs.co put it plainly: "The internet is changing, and you might not have noticed."

That last part is doing a lot of work.

What Pew Actually Measured

The methodology matters here, and it's worth slowing down on it. Pew didn't simply run pages through an AI detector and call it done. According to Digital Trends, the study looked for indicators suggesting AI played a role in producing content — a more honest framing than claiming perfect detection. .com domains showed the biggest increase since ChatGPT launched, per the same report, which tracks: commercial domains are where the economic incentive to automate is sharpest.

The full picture from Pew's own report is that the one-third figure applies specifically to pages published after ChatGPT's release. When you fold in the older web — all those pages written before November 2022 — the share drops considerably. That's not a caveat to dismiss; it's the actual story. The pre-ChatGPT internet is a different substrate. The new internet, the one being built right now, is where the one-third number lives. And per Pew's July 2026 snapshot cited in ForkLog, the trend line is still moving upward.

So: more than a third of everything being added to the web right now shows signs of machine authorship. That's the baseline we're working from.

The Content Farm Déjà Vu Is Not Accidental

Here's my call, since Vincent keeps telling me to stop hedging: this is primarily an acceleration of an existing race to the bottom, not a new threat to quality content that was otherwise flourishing.

The content farm era didn't end because the internet developed better values. It ended because Google changed its algorithm. The underlying incentive — produce as much content as possible for as little cost as possible, capture search traffic, monetize eyeballs — never went away. It just went dormant waiting for cheaper production tools. Those tools arrived in November 2022.

Thin content and spam-optimized pages predate ChatGPT by more than a decade, as SEMrush documents in its breakdown of how thin content proliferated long before generative AI entered the picture. What's changed is the cost curve. Demand Media needed human contractors, however poorly compensated. Today's equivalent operation needs a prompt and an API key. The economics didn't just improve — they collapsed in favor of volume production. The people flooding the zone with AI copy aren't primarily replacing thoughtful human writers; they're replacing the content farm model with something faster and cheaper. The thoughtful human writers were already being squeezed out in 2011.

Where I want to be precise: that argument does not make the Pew finding benign. Scale changes things even when the underlying dynamic is familiar. One-third of new web pages is an enormous surface area. If even a fraction of that content contains errors, outdated information, or subtly biased framing — and there's no reason to assume it won't — the aggregate effect on how people find and trust information is genuinely hard to model. The content farm era was bad for search quality. This could be bad for epistemic quality, which is a different and arguably worse problem.

The Detection Problem Nobody Wants to Talk About

The Pew methodology — looking for indicators of AI authorship rather than claiming definitive detection — is intellectually honest in a way that most of the discourse around this study isn't. Because the follow-on question, the one that matters for anyone trying to do something about this, is: what do you do once you've identified AI-written content?

The detection field is widely debated among researchers, with conflicting evidence about accuracy depending on the model, the domain, and the human editing applied afterward. Pew is explicit that it was looking for signals, not smoking guns. That's the right approach — but it also means the one-third figure should be read as a minimum, a floor, not a ceiling. Content that shows no detectable signs of AI involvement might still have had significant AI involvement; it just had a more careful editor.

Slashdot flagged the senior data scientist framing from Pew in its coverage — the language around pages that "show significant signs" of AI authorship is doing precision work. Significant signs. Not all AI-assisted content rises to that threshold, which means Pew's one-third is almost certainly an undercount of the actual AI contribution to new web content.

What This Actually Demands

The regulatory conversation tends to stall on disclosure — require AI-generated content to be labeled, problem solved. That instinct is understandable and mostly useless. Labeling works when there's a clear production chain. "This page was 60% written by a human who then ran it through Claude and edited the output" does not map cleanly onto a disclosure checkbox.

What the Pew data actually points toward is a search and information quality problem, not primarily an authorship disclosure problem. Google built Panda to clean up the last content farm explosion. The question worth watching is whether the current generation of search and AI answer engines has equivalent leverage — or whether the volume and sophistication of AI-generated content has already outpaced their ability to filter for quality.

The content farm operators of 2009 had a ceiling. They needed humans, even cheap ones, and humans are slow. That ceiling is gone. The Pew trendline going into late 2026 suggests we haven't hit a new one yet.

Whatever the internet looks like on the other side of this, it's being written right now — mostly by machines, mostly for algorithms, mostly in the hope that you won't notice the difference. The content farm era taught us that a lot of people don't. The question is whether the ones who do still have somewhere to go.


Mike Sullivan covers the technology industry for BuzzRAG.

More Like This

Laptop displaying Unreal Engine 5.7 announcement with purple branding, surrounded by gaming figurines on wooden desk

Can Unreal Engine 5 Run on a $500 MacBook? Sort Of.

Testing Unreal Engine 5.7 on the MacBook Neo reveals what happens when professional software meets budget hardware—and why friction matters.

Mike Sullivan·4 months ago·5 min read
Black HDMI 2.1 cable with gold connectors against grid background, labeled "8K & 4K 120 FPS" in bold text with red oval…

Do You Really Need an $80 HDMI Cable? Maybe Not

Tech reviewer Adam tests a premium HDMI 2.1 cable. We examine what you're actually paying for and whether most users need it.

Mike Sullivan·6 months ago·6 min read
A man with long dark hair and a beard speaks on stage at a tech demo day, with "CopilotKit" branding visible and yellow…

When Agents Generate Their Own UI: The Three Flavors Explained

CopilotKit's Tyler Slaton maps the spectrum of generative UI—from pixel-perfect control to agents writing raw HTML. Each approach makes different tradeoffs.

Mike Sullivan·4 months ago·6 min read
MacBook laptop displayed with Unreal Engine logo and Apple M4 chip branding on wooden desk setup

Unreal Engine 5 Still Doesn't Play Nice With Apple Silicon

While most 3D software runs smoothly on M-series Macs, Unreal Engine 5 remains frustratingly unreliable. One creator documents the disconnect.

Mike Sullivan·7 months ago·6 min read
Netflix Used AI in 300 Titles. Here's What That Means.

Netflix Used AI in 300 Titles. Here's What That Means.

Netflix disclosed AI use in roughly 300 titles in its Q2 2026 earnings. We break down what that number actually tells us — and what it doesn't.

Zara Chen·1 month ago·6 min read
Three YouTube channel layouts displaying medieval and historical content with faceless silhouettes, featuring retro pixel…

AI Built a Complete YouTube Video. Here's What It Got Wrong.

A creator typed five prompts. AI researched, scripted, animated, and edited a full video. The logistics worked. The storytelling didn't. Here's what that division means.

Bob Reynolds·1 week ago·6 min read
Man in glasses holding computer parts at a landfill with "NEW SERVER, OLD PARTS" text overlay

When Your Server Dies and Supply Chains Don't Care

Small business sysadmins face a brutal reality: servers die on their own schedule, not the supply chain's. Here's what DIY looks like in 2025.

Mike Sullivan·3 months ago·7 min read
Man in green shirt with stock market charts behind him, text overlay claiming printf is a secret virtual machine

printf: The Tiny Virtual Machine Hiding in Plain Sight

printf isn't just a print function—it's a formatting engine, a security hole, and a tiny VM. Here's what most C programmers never bother to learn about it.

Mike Sullivan·3 months ago·3 min read

RAG·vector embedding

2026-08-25
1,642 tokens1536-dimmodel text-embedding-3-small

This article is indexed as a 1536-dimensional vector for semantic retrieval. Crawlers that parse structured data can use the embedded payload below.