MARATTO

article · International Journal of Computational Intelligence Systems

A Review of Deep Learning-Based Question Generation: Challenges, Methods, and Future Directions

In plain language

Deep learning-based question generation has emerged as a key area within natural language processing, offering applications across education, conversational artificial intelligence, and information retrieval. This review examines the progression of question generation from early rule-based and statistical techniques to modern transformer architectures such as BERT, GPT, T5, and BART. It analyses benchmark datasets and evaluation measures across lexical, semantic, and human-centred dimensions, while also covering the integration of reinforcement learning and multimodal approaches. The review outlines current challenges, including semantic coherence, hallucination, algorithmic bias, evaluation shortfalls, and domain dependence. It highlights promising paths forward such as zero-shot learning, domain adaptation, and human-in-the-loop training to support more robust, context-aware, and ethically sound generation systems.

Key takeaways

  • Question generation has transitioned from rule-based models to modern transformer architectures such as BERT, GPT, T5, and BART.
  • Recent systems incorporate reinforcement learning and multimodal features to improve generation performance.
  • Major persistent challenges include hallucination, bias, domain dependence, and flawed evaluation metrics.
  • Future research prioritises zero-shot learning, domain adaptation, and human-in-the-loop training to improve system robustness.

Why it matters

Automated question generation underpins interactive digital learning tools, smart search systems, and conversational agents. Synthesising current deep learning techniques and their limitations helps engineers and educators build reliable applications that minimise hallucinated or biased content, ultimately ensuring safer and more context-aware interactions for end users.

Commercialisation angle

The review points to potential uses in educational technology, conversational artificial intelligence, and information retrieval tools. Intended users include digital education providers, chatbot developers, and search engine engineers. As this work is an analytical review rather than an applied tool, the underlying technology sits at an exploratory research stage, with commercial deployment reliant on solving identified obstacles such as hallucination and domain dependence.

AI-generated from the published abstract. Always read the original work before citing.

Abstract

Deep Learning-based Question Generation (QG) has become an essential subfield in natural language processing (NLP), with transformative implications across education, conversational AI, and information retrieval. This review critically surveys the evolution of QG techniques from early rule-based and statistical models to modern neural architecture such as sequence-to-sequence models, attention-based frameworks, and transformer-based systems like BERT, GPT, T5, and BART. Key contributions of this paper include a comparative analysis of traditional and deep learning-based approaches, a comprehensive overview of datasets and benchmarks, and a discussion of evaluation metrics encompassing lexical, semantic, and human-centered dimensions. The review further explores the integration of reinforcement learning and multimodal capabilities in state-of-the-art QG systems. It identifies pressing challenges such as semantic coherence, hallucination, bias, evaluation limitations, and domain dependence. By synthesizing recent advances and highlighting future research directions such as zero-shot learning, domain adaptation, and human-in-the-loop training this paper aims to guide researchers and practitioners in developing more robust, adaptable, and ethically sound QG systems. The work serves not only as a resource for understanding current capabilities and limitations but also as a roadmap for advancing the field toward more context-aware, customizable, and high-utility applications.

Research topics

  • Topic Modeling
  • Multimodal Machine Learning Applications
  • Domain Adaptation and Few-Shot Learning

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1007/s44196-026-01497-4

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.