conference paper
Despite the widespread use of Arabic across the internet, the language remains comparatively under-resourced regarding freely accessible, annotated text corpora. To address this deficiency, the Open Source International Arabic News corpus provides a large-scale collection gathered from openly accessible international news platforms. This initial release comprises approximately 3.5 million articles, spanning more than 37 million sentences and roughly one billion tokens. Every article within the collection includes metadata and is formatted in XML. Furthermore, linguistic processing has been applied across the entire text, providing lemma and part-of-speech annotations for every individual word. The completed dataset has been processed, permanently archived, and published within the CLARIN research infrastructure, establishing an openly available linguistic foundation for Arabic language studies and text analysis.
Arabic is widely spoken and read online, yet natural language processing tools and researchers frequently face a shortage of open, structured training data. By providing one billion tokens of linguistically annotated news text within an accessible research infrastructure, this resource directly addresses data scarcity, helping computational linguists and developers study Arabic grammar, improve language models, and build text processing tools.
The resource provides foundational linguistic training data that could enable software developers and technology firms to train Arabic natural language processing tools, such as automated translation, search indexing, or text analysis systems. Because the dataset is already curated, annotated, and hosted on the CLARIN infrastructure, it is ready for immediate deployment in data pipelines, though downstream commercial software built upon it would remain at an early development stage.
AI-generated from the published abstract. Always read the original work before citing.
The World Wide Web has become a fundamental resource for building large text corpora. Broadcasting platforms such as news websites are rich sources of data regarding diverse topics and form a valuable foundation for research. The Arabic language is extensively utilized on the Web. Still, Arabic is relatively an under-resourced language in terms of availability of freely annotated corpora. This paper presents the first version of the Open Source International Arabic News (OSIAN) corpus. The corpus data was collected from international Arabic news websites, all being freely available on the Web. The corpus consists of about 3.5 million articles comprising more than 37 million sentences and roughly 1 billion tokens. It is encoded in XML; each article is annotated with metadata information. Moreover, each word is annotated with lemma and part-ofspeech. The described corpus is processed, archived and published into the CLARIN infrastructure.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.18653/v1/w19-4619
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.