article
Transformer models have pushed forward medical language processing, but most of the research is still mostly in English. That makes it harder to apply these tools effectively in healthcare settings where multiple languages are used. This study presents a twofold contribution. First, we conduct a structured literature review across PubMed, analyzing 1,498 articles and extracting 26 relevant transformer-based models used in biomedical NLP. We give an overview of these models based on language, what they do, and how well they perform. This step helps us better see what's available now in terms of multilingual resources for Arabic medical NLP and where there are still gaps. Second, we construct a new Arabic medical question-answering (QA) dataset comprising over 800,000 QA pairs scraped from public health platforms. A representative sample of 2,500 QA pairs is used to fine-tune and compare two representative models: AraBERTv2 (Arabic-specific) and XLM-RoBERTa (multilingual). The objective of this empirical evaluation is to establish a first baseline on messy, real-world data, rather than to achieve state-of-the-art performance Despite identical training conditions, both models obtain relatively low scores (F1 40 %) which we analyze in light of previous benchmarks and the nature of the data. Our results show that using the latest models on messy, real-world Arabic medical texts is pretty tough. They also point out how important it is to develop better resources customized for less-represented languages like Arabic. Despite identical training conditions, both models obtain relatively low scores (<tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">$\mathbf{F 1}=\mathbf{4 0 \%}$</tex>) which we analyze in light of previous benchmarks and the nature of the data. Our results show that using the latest models on messy, real-world Arabic medical texts is pretty tough. They also point out how important it is to develop better resources customized for less-represented languages like Arabic.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/sita67914.2025.11273428
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.