dataset · Zenodo (CERN European Organization for Nuclear Research)
SARCSenti is a tone-marked Yoruba dataset created to advance affective natural language processing. It is designed for binary sarcasm detection, distinguishing between sarcastic and non-sarcastic content, as well as ternary sentiment classification categorised into negative, neutral, and positive sentiments. The analytical release consists of 1,507 Yoruba headline records formatted in UTF-8 and normalised using Unicode NFC. The data comprises 1,334 records drawn from BBC Yoruba alongside 173 Yoruba translations derived from the Misra Kaggle News Headlines Dataset. Full record-level provenance is documented. Three original records were removed from this analytical release due to missing or invalid sentiment labels. Full-text access is currently restricted owing to unconfirmed redistribution rights for the BBC Yoruba text, though research access remains subject to copyright terms and conditions.
African languages are often under-represented in natural language processing resources. By supplying tone-marked Yoruba text labeled for sentiment and sarcasm, this resource assists developers and researchers in building language tools that accurately capture nuance, emotion, and tonality in low-resource African languages.
The dataset provides early-stage training data that could enable software developers to build natural language processing tools, such as automated sentiment analysis and sarcasm detection systems for Yoruba text. However, commercial application is currently hindered because full-text access is restricted due to unconfirmed copyright and open redistribution rights for the underlying BBC content.
AI-generated from the published abstract. Always read the original work before citing.
SARCSenti is a tone-marked Yoruba dataset developed for affective natural language processing, specifically binary sarcasm detection and ternary sentiment classification. The analytical release contains 1,507 Yoruba headline records, annotated for both sarcasm and sentiment. Sarcasm labels are non-sarcastic (0) and sarcastic (1), while sentiment labels are negative (0), neutral (1), and positive (2). The dataset contains 1,334 BBC Yoruba-derived records and 173 Yoruba translations derived from the Misra/Kaggle News Headlines Dataset for Sarcasm Detection. Record-level provenance is provided in the source field. Three records from the original research workbook were excluded from the analytical release because of missing or invalid numeric sentiment labels. SARCSenti was developed as part of research conducted in the Department of Computer Science, Lead City University, Ibadan, Nigeria. The dataset is intended to support research on Yoruba and African low-resource NLP, sarcasm detection, sentiment analysis, affective computing, cross-task learning, and the processing of tonal languages. Yoruba text is provided in UTF-8 and normalized using Unicode NFC. The accompanying README and data dictionary provide information on dataset structure, labels, provenance, and usage. Access to the full-text dataset is restricted because a substantial portion contains source-derived BBC Yoruba headline text for which open redistribution rights have not yet been confirmed. Research access may be considered subject to applicable source terms and copyright restrictions.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.22308684
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.