MARATTO

dataset · Mendeley Data

QSU: Quran Semantic Units Dataset

Abstract

The Quran Semantic Units (QSU) Dataset, a novel segmentation of the Quranic text that achieves fine-grained granularity by applying the classical rules of Waqf wa Ibtida' (stopping and resuming), which is considered an essential tool for understanding and correctly interpreting the Quranic text. Unlike standard verse-by-verse divisions, which often contain multiple ideas within a single ayah, our dataset partitions the text into self-contained, semantically complete units. Built to advance NLP applications such as semantic search, question-answering, contextual embeddings, and automated tafsir, our methodology synthesizes established stopping conventions with rigorous linguistic analysis to accurately delineate semantic boundaries.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.17632/v7hyhk7krd

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.