article
In recent years, the Vision-Language Models (VLMs) have significantly advanced the field of storytelling applications. Visual storytelling involves the use of Generative AI models to transform a sequence of video frames into a semantics-preserving story-like text. A critical aspect of developing these models lies in the preparation of datasets, which directly impacts their efficacy and robustness. This paper explores a comprehensive methodology for dataset preparation tailored to the detection and description of car accidents in video surveillance footage. This research explains the necessity of data collection, annotation, and preprocessing to transform a sequence of frames in a video into a semantics-preserving story-like text. We describe the steps, tools and challenges of this task.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iceeng64546.2025.11031376
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.