How AI Training Data Gets Collected, and Why It's Controversial
AI training data collection has become a major point of controversy. Here's how it actually happens.
Training a large AI model typically requires enormous amounts of text, images, or other data, much of which has historically been gathered by automatically scanning publicly accessible parts of the internet.
Why this collection method is controversial
Much of this scanned content was originally created by individual writers, artists, and photographers who never explicitly agreed to have their work used to train a commercial AI system, leading to significant ongoing legal disputes over whether this widespread practice requires consent or compensation.
How this is starting to change
Some AI companies have begun signing licensing agreements directly with publishers and content platforms, and some countries have introduced or proposed regulations requiring greater transparency about what data was actually used to train a given model, both responses to the ongoing controversy over how training data has traditionally been collected.