MARATTO

article

Topic Modelling Swahili Using LDA and Contextualized Embeddings

Abstract

Technological innovation in today's society has led to a highly interconnected world, the fruit of which has been an unprecedented breakthrough in new technologies and a tremendous growth of information resources. The availability of online platforms with user-generated content has increased the volume of textual data, leading to a great need for better tools and methods to process the data into meaningful information. While there has been a lot of research on topic modelling in high-resource languages, there is limited research in resource-scarce languages like Swahili despite being spoken by millions world-wide. New topic modelling techniques provide an opportunity to advance NLP research in low-resource settings. This paper provides benchmark experiments for Swahili topic modelling using different transformer based pre-trained language models as embeddings and compares the results with the baseline LDA technique. We also conduct experiments to examine the impact of two dimensionality reduction techniques, PCA and UMAP, on the performance of the topic modelling framework in long Swahili texts. We evaluate the resulting topic models using the <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$C\_V$</tex> and Normalized pointwise Mutual information (npmi) coherence scores. The results show that BERTopic using XLM-Roberta Pretrained Language model embeddings with UMAP dimensionality reduction and HDBSCAN clustering was best suited for topic modelling Swahili, achieving a c_v coherence of 0.776 and nmpi of 0.199. This was followed by the model developed using SwahBERT embeddings with UMAP and HDBSCAN clustering, which achieved a c_v Coherence of 0.707 and npmi of 0.171. Our findings provide a benchmark NLP contribution in automatically categorising Swahili news articles using topic modelling techniques.

Research topics

  • Topic Modeling
  • Advanced Text Analysis Techniques
  • Sentiment Analysis and Opinion Mining

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/imsa61967.2024.10652736

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.