Newly unsealed legal documents have revealed that a Microsoft executive strongly criticized the web scraping methods used to train AI models, describing them internally using the harsh phrase "the largest labor exploitation in human history." This revelation came to light amid an ongoing series of legal disputes surrounding copyright and the ethics of data collection in AI development.
This matter is not an announcement regarding a specific product update; rather, it indicates that deep ethical concerns were debated internally within the company regarding the data collection processes that serve as the foundation of Microsoft's AI development. While the company has previously maintained the legality of its data collection, this description in internal documents highlights a disconnect within the organization regarding transparency and rights protection in Large Language Model (LLM) development.
The acquisition of massive training datasets—including through the partnership between Microsoft and OpenAI—continues to face persistent criticism from creators and the media industry. This disclosure has the potential to accelerate future legal regulations and demands for stricter transparency regarding the data sources used in AI models. Companies will be expected to address the difficult challenge of establishing legal and ethical data collection methods while maintaining technological innovation.