If we compare AI large models to cars, the raw data is like crude oil.
Author: Jiang Jiang
Editor: Manman Zhou
The emergence of ChatGPT and the explosive adoption of Midjourney have enabled AI to achieve its first large-scale application, that is, the popularization of large models.
Large models refer to machine learning models with a large number of parameters and complex structures, capable of processing massive data and performing various complex tasks.
01 AI Data Copyright Disputes
If we compare today's AI large models to cars, the raw data is like crude oil. In any case, first and foremost, AI models need sufficient "crude oil."
The "crude oil" sources of AI companies mainly include the following:
- Publicly available online data sources, such as Wikipedia, blogs, forums, news, etc.
- Traditional news media and publishers
- Universities and research institutions
- C-end users of the models
The ownership of real-world oil has mature legal regulations, but in the chaotic field of AI, the ownership of "crude oil" extraction rights is not clear, leading to numerous disputes.
Recently, several major music labels sued AI music production companies Suno and Udio, accusing them of copyright infringement. This lawsuit is similar to the one filed by The New York Times against OpenAI in December last year.

Image source: Billboard
In July 2023, some authors filed a lawsuit against the company, accusing ChatGPT of generating summaries of authors' works based on copyrighted content.
In December of the same year, The New York Times also filed a similar copyright infringement lawsuit against Microsoft and OpenAI, accusing these two companies of using the newspaper's content to train artificial intelligence chatbots.
Additionally, a class-action lawsuit was filed in California, accusing OpenAI of obtaining users' private information from the internet without consent to train ChatGPT.
Ultimately, OpenAI did not bear the brunt of these accusations. They stated that they did not agree with The New York Times' accusations, could not reproduce the issues mentioned by The New York Times, and, more importantly, the data sources provided by The New York Times were not important to OpenAI.

Source: https://openai.com/index/openai-and-journalism/
For OpenAI, perhaps the biggest lesson from this incident is to manage the relationship with data suppliers and clarify the rights and responsibilities of both parties. As a result, over the past year, we have seen OpenAI enter into partnerships with many data suppliers, including but not limited to The Atlantic, Vox Media, News Corp, Reddit, Financial Times, Le Monde, Prisa Media, Axel Springer, American Journalism Project, and more.
In the future, OpenAI will legitimately use the data from these media outlets, and these media outlets will also integrate OpenAI's technology into their products.
02 AI Driving Content Platform Monetization
However, the fundamental reason for OpenAI's partnership with data suppliers is not the fear of being sued, but rather the impending data depletion faced by machine learning. Researchers at MIT and others have estimated that machine learning datasets may deplete all "high-quality language data" by 2026.
As a result, "high-quality data" has become a hot commodity for model manufacturers like OpenAI and Google. Content companies have repeatedly entered into partnerships with AI model manufacturers, initiating a passive income mode.
Traditional media platform Shutterstock has successively partnered with Meta, Alphabet, Amazon, Apple, OpenAI, Reka, and other AI companies. In 2023, the revenue from content licensing to AI models increased to $104 million, and it is expected to generate $250 million in revenue by 2027. Reddit's content copyright revenue licensed to Google reaches as high as $60 million annually. Apple is also seeking to collaborate with mainstream news media, offering at least $50 million in annual copyright fees. The copyright fees received by content companies from AI companies are skyrocketing at an annual growth rate of 450%.

Image source: CX Scoop
In recent years, content outside of streaming has been difficult to monetize, which is a major pain point for the content industry. Compared to the internet entrepreneurial era, the emergence of AI has brought greater imagination and stronger revenue expectations to the content industry.
03 High-Quality Data Still Scarce
Of course, not all types of content meet the needs of AI.
Regarding the earlier mentioned dispute between OpenAI and The New York Times, another highlight is the data quality. Just as refining oil from crude oil requires good quality oil and good refining technology, OpenAI specifically emphasized that The New York Times' content did not make any significant contribution to the training of OpenAI's models. Compared to Shutterstock, which can cost OpenAI tens of millions of dollars annually, The New York Times, a text media that relies on timeliness, is not the darling of the AI era. AI requires more profound and unique data.
As high-quality data is scarce, AI companies have also begun to focus on "refining technology" and "one-stop applications."
On June 25, OpenAI acquired real-time analytics database company Rockset. This company primarily provides real-time data indexing and querying capabilities, and OpenAI plans to integrate Rockset's technology into its products to enhance the real-time value of data usage.

Image source: DePIN Scan
Through the acquisition of Rockset, OpenAI plans to enable AI to better utilize and access real-time data. This will enable OpenAI's products to support more complex applications, such as real-time recommendation systems, dynamically data-driven chatbots, real-time monitoring and alert systems, and more.
Rocket is OpenAI's internal "petrochemical department," directly converting ordinary data into high-quality data required for applications.
04 Is It Unrealistic for Creators to Assert Data Rights?
The data of internet media platforms (Facebook, Reddit, etc.) largely comes from UGC, i.e., user-generated content. Many platforms, while charging high data fees to AI companies, have quietly added a clause to their terms of service stating that the platform has the right to use user data to train AI models.
Although the terms of service specify the rights to train AI models, many authors are not clear about which models are using the content they produce, whether it is used for payment, and have no way to obtain the relevant rights that should belong to them.
During Meta's quarterly earnings call in February of this year, Zuckerberg explicitly stated that he would use images from Facebook and Instagram to train his AI generation tool.
It has been reported that Tumblr has also reached mysterious content licensing agreements with OpenAI and Midjourney, but the specific content of the agreements has not been publicly disclosed.
The creators on the image platform EyeEm recently received a notification informing them that the photos they have published will be used for AI model training. The notification mentioned that users can choose not to use the product as a result, but did not mention any compensation policy. EyeEm's parent company, Freepik, revealed to Reuters that the company has signed agreements with two large tech companies to license most of the 200 million images for around 3 cents per image. CEO Joaquin Cuenca Abela stated that five similar deals are in progress but refused to disclose the buyers' identities.

UGC-dominated content platforms such as Getty Images, Adobe, Photobucket, Flickr, and Reddit are all facing similar issues. Driven by the huge potential for data monetization, these platforms choose to ignore users' content ownership and sell the data to AI model companies.
The entire process is conducted in secret, and creators have no opportunity to resist. Many creators may not even have the chance to question whether their works were sold to AI companies for model training until one day in the future when similar content is trained in a model.
Web3 may be a good choice to address the issue of protecting creators' data rights and earnings. While AI companies are reaching new highs in the stock market, the concept of AI coins in Web3 is also soaring. With its decentralized and tamper-proof characteristics, blockchain enjoys a unique advantage in protecting creators' rights.
Media content such as images and videos completed large-scale on-chain adoption during the bull market of 2021, and the on-chain adoption of UGC content on social platforms is quietly happening. At the same time, many Web3 AI model platforms are already incentivizing ordinary users who contribute to model training, whether as data owners or trainers.
The exponential development of AI models has created a greater demand for data rights. Creators should consider why their works are being sold to AI model companies for 5 cents per piece without their consent. They should also question why they are unaware of the entire process and unable to receive any benefits.
Exploiting the media platforms to the fullest cannot alleviate the data anxiety of AI model companies. The prerequisite for achieving high-quality data at high volume is data rights, which involves a fair distribution of interests among creators, platforms, and AI model companies.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。