ChatGPT

How AI quietly reshaped a third of the modern web

A third of web pages published since ChatGPT's launch now show signs of AI authorship.

3 min readTechCrunch
How AI quietly reshaped a third of the modern web

A third of the web pages published since ChatGPT's launch show signs of AI authorship. That statistic should stop you, because it is not about the technology's existence. It is about the quiet normalization of machine-written text in spaces we once trusted as human. For anyone who works with data, this is not a curiosity. It is a condition of the environment. The same tools that make spreadsheet formulas conversational are now generating the articles, product descriptions, and forum posts we scrape for training sets. If you are building models on that material, you are not capturing the web. You are capturing a feedback loop.

This is where the story connects to our own work. We recently explored how Clean Data Starts With Catching AI Slop Before It Skews Your Model, and the finding was uncomfortable: filtering out AI-flagged content made sentiment models less accurate because many genuine reviews were being thrown out. That is the trap. The web is not cleanly divided into human and machine. It is a mixture where the proportion of AI text is rising, and any binary filter will fail. The study's headline number is not a judgment about quality. It is a measurement of volume. But volume matters because it changes the baseline. If a third of new pages are AI-assisted, then your training data is already shaped by the very systems you are trying to evaluate. You cannot opt out. You can only get better at recognizing the texture of that influence.

Our take is not that AI authorship is inherently bad. Some of the most useful technical explanations on the web are now drafted with AI assistance, then edited by humans who know their subject. That is a workflow, not a fraud. The real risk is complacency. If you assume that human authorship guarantees reliability, or that machine authorship guarantees sloppiness, you will misread your sources. We have seen this in practice with Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where practical deployment challenges rarely come from the model itself, but from the assumptions baked into the data pipeline. The same is true here. The model is not the problem. The unexamined dataset is.

What would we tell a reader who asks about this? Start treating authorship as a feature, not a bug. When you evaluate a source, ask whether the text was generated, edited, or merely inspired by AI, because that distinction changes how much weight you should give it. The tools we build for spreadsheet analysis are becoming more capable, but they are only as clear as the information we feed them. If a third of the web is now machine-influenced, then your next model is already learning from that reality. The concrete point to watch is not the percentage itself. It is how quickly your own filtering methods adapt. The moment you assume your data is clean, you have already introduced the bias you were trying to avoid.

From TechCrunch

ChatGPT and other AI models are now authoring and editing much of the new web.

Read the original at TechCrunch