AI Ethics & Policy

How AI Training Data Gets Collected, and Why It's Controversial

AI training data collection has become a major point of controversy. Here's how it actually happens.

1 min read · AI & Machine Learning

Training a large AI model typically requires enormous amounts of text, images, or other data, much of which has historically been gathered by automatically scanning publicly accessible parts of the internet.

Why this collection method is controversial

Much of this scanned content was originally created by individual writers, artists, and photographers who never explicitly agreed to have their work used to train a commercial AI system, leading to significant ongoing legal disputes over whether this widespread practice requires consent or compensation.

How this is starting to change

Some AI companies have begun signing licensing agreements directly with publishers and content platforms, and some countries have introduced or proposed regulations requiring greater transparency about what data was actually used to train a given model, both responses to the ongoing controversy over how training data has traditionally been collected.