article · Journal of Applied Learning & Teaching
A systematic review evaluated seventeen studies published in 2023 examining how effectively artificial intelligence detectors distinguish between human-written and machine-generated text. Across the analysed research, ChatGPT models GPT-3.5 and GPT-4 were consistently evaluated. Performance varied considerably across different detection platforms. Crossplag emerged as the most effective detection tool, followed closely by Copyleaks, whereas Duplichecker and Writer delivered the poorest detection outcomes. Crucially, the review identified widespread inconsistency and a pronounced lack of reliability across all evaluated artificial intelligence detection systems and conventional anti-plagiarism tools. Because automated tools alone do not offer dependable results, differentiating synthetic text from human writing currently requires combining contemporary artificial intelligence detectors, standard anti-plagiarism software, and human evaluation.
The growing use of generative text models like ChatGPT creates challenges for verifying authorship in education, publishing, and professional environments. Because current artificial intelligence detection tools deliver unreliable and inconsistent results, organisations cannot depend on fully automated solutions alone. Understanding these limitations helps educators and administrators develop more realistic, hybrid verification practices that combine digital detection tools with human oversight.
Educational institutions, academic publishers, and corporate compliance teams seeking to verify text authenticity represent key user groups for detection technologies. However, commercially available detection software remains inconsistent and unreliable on its own. Rather than deploying standalone automated products, near-term implementation pathways require hybrid verification systems that integrate contemporary artificial intelligence detectors, standard anti-plagiarism utilities, and human reviewers to assess content authenticity effectively.
AI-generated from the published abstract. Always read the original work before citing.
The purpose of this study was to review 17 articles published between January 2023 and November 2023 that dealt with the performance of AI detectors in differentiating between AI-generated and human-written texts. Employing a slightly modified version of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) protocol and an aggregated set of quality evaluation criteria adapted from A MeaSurement Tool to Assess systematic Reviews (AMSTAR) tool, the study was conducted from 1 October 2023 to 30 November 2023 and guided by six research questions. The study conducted its searches on eleven online databases, two Internet search engines, and one academic social networking site. The geolocation and authorship of the 17 reviewed articles were spread across twelve countries in both the Global North and the Global South. ChatGPT (in its two versions, GPT-3.5 and GPT-4) was the sole AI text generator used or was one of the AI text generators in instances where more than one AI text generator had been used. Crossplag was the top-performing AI detection tool, followed by Copyleaks. Duplichecker and Writer were the worst-performing AI detection tools in instances in which they had been used. One of the major aspects flagged by the main findings of the 17 reviewed articles is the inconsistency of the detection efficacy of all the tested AI detectors and all the tested anti-plagiarism detection tools. Both sets of detection tools were found to lack detection reliability. As a result, this study recommends utilising both contemporary AI detectors and traditional anti-plagiarism detection tools, together with human reviewers/raters, in an ongoing search for differentiating between AI-generated and human-written texts.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.37074/jalt.2024.7.1.14
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.