MARATTO

article

Dataset Preprocessing Leveraging Vision Language Models: Storytelling for Accident Detection and Description

Abstract

In recent years, the Vision-Language Models (VLMs) have significantly advanced the field of storytelling applications. Visual storytelling involves the use of Generative AI models to transform a sequence of video frames into a semantics-preserving story-like text. A critical aspect of developing these models lies in the preparation of datasets, which directly impacts their efficacy and robustness. This paper explores a comprehensive methodology for dataset preparation tailored to the detection and description of car accidents in video surveillance footage. This research explains the necessity of data collection, annotation, and preprocessing to transform a sequence of frames in a video into a semantics-preserving story-like text. We describe the steps, tools and challenges of this task.

Research topics

  • Fire Detection and Safety Systems
  • Safety Warnings and Signage
  • IoT and GPS-based Vehicle Safety Systems

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/iceeng64546.2025.11031376

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.