Their findings, shared solely with MIT Know-how Assessment, present a worrying pattern: AI’s information practices threat concentrating energy overwhelmingly within the arms of some dominant expertise corporations.
Within the early 2010s, information units got here from quite a lot of sources, says Shayne Longpre, a researcher at MIT who’s a part of the venture.
It got here not simply from encyclopedias and the online, but additionally from sources similar to parliamentary transcripts, incomes calls, and climate experiences. Again then, AI information units have been particularly curated and picked up from completely different sources to go well with particular person duties, Longpre says.
Then transformers, the structure underpinning language fashions, have been invented in 2017, and the AI sector began seeing efficiency get higher the larger the fashions and information units have been. Right this moment, most AI information units are constructed by indiscriminately hoovering materials from the web. Since 2018, the online has been the dominant supply for information units utilized in all media, similar to audio, photos, and video, and a niche between scraped information and extra curated information units has emerged and widened.
“In basis mannequin growth, nothing appears to matter extra for the capabilities than the dimensions and heterogeneity of the information and the online,” says Longpre. The necessity for scale has additionally boosted using artificial information massively.
The previous few years have additionally seen the rise of multimodal generative AI fashions, which may generate movies and pictures. Like massive language fashions, they want as a lot information as attainable, and the very best supply for that has turn into YouTube.
For video fashions, as you possibly can see on this chart, over 70% of knowledge for each speech and picture information units comes from one supply.
This might be a boon for Alphabet, Google’s father or mother firm, which owns YouTube. Whereas textual content is distributed throughout the online and managed by many various web sites and platforms, video information is extraordinarily concentrated in a single platform.