MARATTO

article

Experimental Insights into Multilingual Language Detection Using Machine and Deep Learning Models with Balanced Training Strategies

Abstract

Language recognition is a fundamental task in natural language processing (NLP) that enables the development of multilingual systems for applications such as information retrieval, machine translation, and text classification. This paper presents a comparative study of machine learning (ML) and deep learning (DL) approaches for automatic language identification across a large corpus of multilingual text. The methodology includes rigorous preprocessing steps, feature extraction using TF-IDF for ML models, and embeddings for DL models. Several classifiers, including Naïve Bayes, Logistic Regression, Support Vector Machines, Decision Trees, Convolutional Neural Networks (CNN), and Recurrent Neural Networks (RNN), are trained and evaluated. Standard performance metrics-accuracy, precision, recall, and F1-score-are used to assess the models. Experimental results show that while ML classifiers achieve competitive baselines, DL architectures consistently outperform them, particularly in handling large and diverse multilingual datasets. These findings highlight the role of contextual representations and neural architectures in building scalable and robust language recognition systems.

Research topics

  • Natural Language Processing Techniques
  • Speech Recognition and Synthesis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/miucc66482.2025.11196745

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.