article · E3S Web of Conferences
Personalized marketing content has become a key driver for customer engagement in modern digital platforms. However, generating customized video content at scale remains a significant challenge due to the need for dynamic adaptation and automation. The proposed pipeline leverages Whisper-based speech transcription and Coqui TTS voice cloning to perform CSV driven keyword replacement, enabling the automatic generation and delivery of one personalized video per client entry. This paper proposes an automated and scalable system for generating personalized marketing videos based on structured client profiles and predefined multimedia templates. The approach integrates client data preprocessing, dynamic content selection, and automated video composition within a unified framework. Experimental validation confirms the feasibility and robustness of the proposed system, demonstrating its capability to efficiently generate customized marketing videos while ensuring scalability and flexibility for real-world deployment.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1051/e3sconf/202669801016
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.