MARATTO

article

Unraveling the Parameters of K-Means Clustering for Enhanced Textual Analysis

Abstract

Text classification is a crucial operation for the management of voluminous textual data, the sheer quantity of data in different establishments (companies - universities…) have made their classification into homogeneous and meaningful clusters paramount and to perform this task the K-means algorithm remains the most widely used choice for its low resource consumption and remarkably accurate results. In this study we attempt to explore the parameter configurations of the K-means method by analyzing two distinct text corpora: one from sports articles and the other from news articles. Through meticulous experimentation, we investigate the impact of text pre-processing, similarity measures and pre-trained linguistic models in order to specify the best configuration for using K-means in text classification. The results underline the importance of pre-processing, identify cosine similarity as optimal and highlight the potential of pre-trained models. Our insights provide concrete recommendations for practitioners, advancing text clustering methodology for a variety of applications.

Research topics

  • Text and Document Classification Technologies

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/unet62310.2024.10794714

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.