MARATTO

article · International Journal of Computing

Referencing of Document Content Using Similarity Measures

Abstract

One of the biggest challenges with scientific writing automation is still the difficulty of automatically locating and adding relevant references in scholarly papers. This paper addresses this issue by proposing a three-phase automatic referencing system based on semantic similarity measures: reference insertion, semantic similarity computation, and preprocessing (tokenization, stop word removal, morphosyntactic marking, and lemmatization). Based on semantic similarity, our experimental results confirm that the system can automatically identify and insert relevant references. The Resnik measure outperformed the Mihalcea measure (43% accuracy, 50% precision, and 59% F1-score), achieving the best performance with (57% accuracy, 58% precision, and 64% F1-score).

Research topics

  • Natural Language Processing Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.47839/ijc.24.2.4010

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.