IAD Index of Academic Documents
  • Home Page
  • About
    • About Izmir Academy Association
    • About IAD Index
    • IAD Team
    • IAD Logos and Links
    • Policies
    • Contact
  • Submit A Journal
  • Submit A Conference
  • Submit Paper/Book
    • Submit a Preprint
    • Submit a Book
  • Contact
  • Dokuz Eylül Üniversitesi Mühendislik Fakültesi Fen ve Dergisi
  • Volume:27 Issue:79
  • Performance Comparison of Text Weighting Schemas on NMF-Based Topic Analysis

Performance Comparison of Text Weighting Schemas on NMF-Based Topic Analysis

Authors : Tolga Berber, Melek Eriş Büyükkaya
Pages : 46-53
Doi:10.21205/deufmd.2025277907
View : 33 | Download : 52
Publication Date : 2025-01-23
Article Type : Research Paper
Abstract :Nowadays, it is feasible to analyze text data that is being generated at an exponential rate by transforming it into a sparse matrix of big size using a certain weighting method. A comprehensive text weighting approach consists of three fundamental components: Term Frequency, Document Frequency, and Vector Normalization. The multiplication of these three components yields numerical values that indicate the significance of a word for a text. Nevertheless, the unprocessed state of these values is unsuitable for the semantic analysis of textual material. There are multiple techniques available for this objective, and Topic Analysis, which seeks to identify subjects discussed in extensive text collections, is one of these techniques. The Non-Negative Matrix Factorization (NMF) approach is commonly employed in topic analysis. It involves transforming an input matrix into the product of two or more matrices, using both random and deterministic beginning values. This study involved conducting tests on a dataset of 20,000 articles sourced from Wikipedia, the online encyclopedia, with the aim of investigating the impact of text weighting methods and initial value approaches commonly employed in the literature on the NMF method. The number of clusters to be used in the studies was determined using an analytical procedure, which employed an upper limit. The results indicate that the “lnc” and “nnc” weighting schemes yielded the highest performance in NMF. These findings demonstrate that employing the “lnc” or “nnc” weighting scheme will lead to more favorable outcomes in the domain of topic analysis.
Keywords : Konu Analizi, Metin Ağırlıklandırma Şemaları, Negatif Olmayan Matris Ayrışımı, Başarım Karşılaştırması

ORIGINAL ARTICLE URL
VIEW PAPER (PDF)

* There may have been changes in the journal, article,conference, book, preprint etc. informations. Therefore, it would be appropriate to follow the information on the official page of the source. The information here is shared for informational purposes. IAD is not responsible for incorrect or missing information.


Index of Academic Documents
İzmir Academy Association
CopyRight © 2023-2025