MARATTO

article

Co-Story: Consistent, Training-Free Story Visualization Using Latent Diffusion and ControlNet

Abstract

Transforming written stories into a sequence of images, known as story visualization, brings narratives to life by combining language with visuals. However, this task poses several challenges, such as maintaining scene and character consistency and aligning images with text. Recent approaches employ diffusion models and large language models (LLMs) to tackle these issues. While progress has been made, especially in preserving character identity, achieving full scene coherence without costly training remains difficult. This paper introduces Latent Context Injection (LCI) module that uses latent encoding of a weighted blend of prior images to enhance scene consistency. To prevent characters from appearing rigid, pose guidance is applied to reflect narrative actions. Additionally, Low-Rank Adaptation (LoRA) is used to efficiently capture and retain the distinctive appearance of new characters with minimal fine-tuning, thereby improving scalability, visual fidelity, and scene consistency.

Research topics

  • Human Motion and Animation
  • Video Analysis and Summarization
  • Human Pose and Action Recognition

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/miucc66482.2025.11196824

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.