Библиографическое описание:System of Automated Text Messages Clustering by Semantic Proximity Based on NLP and Machine Learning Methods : доклад, тезисы доклада / I. I. Khudonogova, L. L. Lipinskiy, A. A. Polyakova. - [S. l. : s. n.], 2023. - Текст : непосредственный // Hybrid methods of modeling and optimization in complex systems : Proceedings of the International Workshop “Hybrid methods of modeling and optimization in complex systems” (HMMOCS 2022) / International Workshop “Hybrid methods of modeling and optimization in complex systems” (HMMOCS 2022) (2022 ; 22.11 - 24.11 ; Krasnoyarsk). - London, United Kingdom, 2023. - P. 19-31. - ISBN 9781802969603, DOI 10.15405/epct.23021.3.
Аннотация:At the present moment the relevance of natural data processing problem solving is rising. A massive data amount of text data has been accumulated in recent years. Classical analytical methods, such as machine learning methods, are not capable of dealing with raw text data, which complicates the analysis significantly. Therefore, a modern set of methods of text data vectorization has been developed, which gained massive popularity in the recent years for analyzing text data, specifically for solving text clustering problem, as one of the most relevant text data related analytical problems. In this paper, a few of these methods were researched; a new dictionary optimization approach has been proposed and tested on the real text datasets; a number of conclusions on the effectiveness and of the methods for the given tasks has been made. For the future work a more thorough research on the dictionary optimization scheme (genetic algorithm parameters) and vectorization method are planned. ]