Kaznewsdataset: Single country overall digital mass media publication corpus

Yakunin K. Kalimoldayev M. Mukhamediev R.I. Mussabayev R. Barakhnin V. Kuchin Y. Murzakhmetov S. Buldybayev T. Ospanova U. Yelis M. Zhumabayev A. Gopejenko V. Meirambekkyzy Z. Abdurazakov A.
March 2021 MDPI

Data
2021 #6 Issue 3

Mass media is one of the most important elements influencing the information environment of society. The mass media is not only a source of information about what is happening but is often the authority that shapes the information agenda, the boundaries, and forms of discussion on socially relevant topics. A multifaceted and, where possible, quantitative assessment of mass media performance is crucial for understanding their objectivity, tone, thematic focus and, quality. The paper presents a corpus of Kazakhstan media, which contains over 4 million publications from 36 primary sources (which has at least 500 publications). The corpus also includes more than 2 million texts of Russian media for comparative analysis of publication activity of the countries, also about 4000 sections of state policy documents. The paper briefly describes the natural language processing and multiple-criteria decision-making methods, which are the algorithmic basis of the text and mass media evaluation method, and describes the results of several research cases, such as identification of propaganda, assessment of the tone of publications, calculation of the level of socially relevant negativity, comparative analysis of publication activity in the field of renewable energy. Experiments confirm the general possibility of evaluating the socially significant news, identifying texts with propagandistic content, evaluating the sentiment of publications using the topic model of the text corpus since the area under receiver operating characteristics curve (ROC AUC) values of 0.81, 0.73 and 0.93 were achieved on abovementioned tasks. The described cases do not exhaust the possibilities of thematic, tonal, dynamic, etc., analysis of the considered corpus of texts. The corpus will be interesting to researchers considering both multiple publications and mass media analysis, including comparative analysis and identification of common patterns inherent in the media of different countries.

ARTM , Computer modeling , LDA , Mass-media , Multiple-criteria decision-making (MCDM) , Natural language processing , Propaganda identification , Sentiment analysis , Significant social news , Topic modeling

Text of the article Перейти на текст статьи

Institute of Information and Computational Technologies, Almaty, 050010, Kazakhstan
Institute of Cybernetics and Information Technology, Satbayev University (KazNRTU), Almaty, 050013, Kazakhstan
Department of Natural Science and Computer Technologies, ISMA University, Riga, LV-1011, Latvia
Federal Research Center for Information and Computational Technologies, Novosibirsk, 630090, Russian Federation
Department of Information Technologies, Novosibirsk State University, Novosibirsk, 630090, Russian Federation
Information-Analytical Center, Nur-Sultan, 010000, Kazakhstan
International Radio Astronomy Centre, Ventspils University of Applied Sciences, Ventspils, LV-3601, Latvia

Institute of Information and Computational Technologies
Institute of Cybernetics and Information Technology
Department of Natural Science and Computer Technologies
Federal Research Center for Information and Computational Technologies
Department of Information Technologies
Information-Analytical Center
International Radio Astronomy Centre

10 лет помогаем публиковать статьи Международный издатель

Книга Публикация научной статьи Волощук 2026 Book Publication of a scientific article 2026