-
Media Cloud: Massive Open Source Collection of Global News on the Open Web
Abstract: We present the first full description of Media Cloud, an open source platform based on crawling hyperlink structure in operation for over 10 years, that for many uses will be the best way to collect data for studying the media ecosystem on the open web. We document the key choices behind what data Media Cloud collects and stores, how it processes and organizes these data, and its open API access a… ▽ More
Submitted 1 May, 2021; v1 submitted 8 April, 2021; originally announced April 2021.
Comments: 15 pages, 9 figures, accepted (minus the 3-page, 3-image appendix given here) for publication and forthcoming in Proceedings of the Fifteenth International AAAI Conference on Web and Social Media (ICWSM-2021)
ACM Class: J.4; H.3.5; J.7; J.5; K.4.1