article
Information extraction is used to automatically extract useful information from unstructured or semi-structured data. The exponential rise of unstructured, multidimensional data poses novel challenges for information extraction techniques in the era of big data. The volume and variety of big data imply the need to implement advanced techniques to extract information reliably and at scale. This paper introduces a novel approach for distributed data extraction based on the combined use of PaddleOCR and the distributed framework Apache Kafka. This work aims to present and compare four distinct methods for deploying PaddleOCR in a distributed environment, thereby integrating AI-enhanced models into a scalable and resilient framework. This article provides a roadmap for developing large-scale information extraction systems that are both high-performing and adaptable, meeting the growing demands for data processing in ever-evolving digital environments.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/niss66502.2025.00027
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.