AI models have spent years learning from the internet. But what happens when the internet itself becomes increasingly filled with content created by AI?
Researchers have a name for one potential consequence: model collapse.
The idea is relatively simple. Large AI models are trained on enormous collections of human-created text, images, code and other information. But generative AI is rapidly changing that information environment. The internet is increasingly populated with AI-written articles, synthetic images, generated code, automated reviews and machine-produced social media posts.
That means future AI models may increasingly learn not just from us, but from the outputs of previous AI models.
How Does It Happen?
Imagine the first generation of an AI system being trained predominantly on original human-created information. That model then generates billions of pieces of new content, some of which find their way onto websites, social media platforms and other parts of the internet.
A future model collects information from that same internet, except this time its training data contains some of the material produced by the previous generation of AI. That model creates even more synthetic content, which becomes part of the information environment encountered by the generation after that.
The cycle can repeat: humans create data, AI learns from it, AI creates new data, and future AI learns from those outputs.
Researchers studying recursive training have shown that, under certain conditions, repeatedly training models on generated data can distort the underlying distribution they are trying to learn. Less common information can gradually disappear while already-common patterns become increasingly dominant.
This is model collapse.
Think of a Photocopy of a Photocopy
There is a simple way to understand the problem.
Imagine photocopying an original photograph. The first copy might look almost identical to the original. Now photocopy the copy. Then photocopy that copy again.
With every generation, tiny imperfections can accumulate. Fine details disappear, distortions become more noticeable and eventually the copy becomes a poor representation of the original.
Something similar can happen when AI repeatedly learns from synthetic outputs without sufficient access to high-quality original data.
It is tempting to describe this as AI simply becoming “stupider”, but the problem is more subtle. The model’s representation of the world can become narrower, more repetitive and less representative of the original information from which earlier generations learned.
The AI Ouroboros
There is an ancient symbol that captures this surprisingly well: the ouroboros, the snake eating its own tail.
AI initially consumes human knowledge. It then produces enormous quantities of new content. That content enters the information ecosystem, where another generation of AI may eventually consume it again.
The machine begins eating its own informational tail.
And that creates an important question for the future of artificial intelligence: how do we know where the information used to train a model originally came from?
This Is Also a Trust Problem
Model collapse is not simply a technical question about whether a benchmark score rises or falls. It is potentially a question of trust.
Businesses, schools, universities, governments and other organisations are increasingly using AI systems for research, software development, analysis, communication and decision support. Reliability therefore matters.
If organisations cannot establish the provenance and quality of the information used to develop an AI system, it becomes harder to understand the limitations of what that system has learned.
Was the underlying information created by a person? Was it generated by another AI? Was that AI output itself based partly on previous synthetic information? Were mistakes introduced or amplified somewhere along the way?
As these chains become longer, data provenance may become just as important as data volume.
Human Data Could Become Premium
There is an interesting twist to all of this.
The more synthetic content we create, the more valuable high-quality human-generated information may become.
Original scientific research, expert writing, specialist knowledge, carefully curated datasets, real human conversations and observations of the physical world could become increasingly important resources for training and evaluating future AI systems.
The next race in artificial intelligence may therefore not simply be about who has the biggest model or the most computing power. It could also be about who has access to trustworthy, original and verifiable information.
That does not mean synthetic data is inherently bad. Carefully designed synthetic data can be extremely useful, and researchers are developing techniques for using it effectively.
The problem emerges when synthetic information becomes an unquestioned replacement for the diverse human information from which these systems originally learned.
The Bigger Picture
For years, much of the public debate about artificial intelligence has focused on one fear: what happens if AI becomes too powerful?
Perhaps there is another question worth asking.
What happens if AI becomes so common that it begins changing the very information environment future AI depends upon?
The internet was once predominantly a record of human activity. Increasingly, humans and machines are writing that record together.
That means one of the defining challenges of the next generation of AI may not simply be building systems capable of learning more.
It may be proving where their knowledge came from and whether we can still trust it.
We spent years worrying about AI becoming too powerful. Perhaps we should also worry about AI becoming so common that it starts learning from itself.
