MARATTO

article

AI-Enhanced Data Extraction Models for Scalabale and Efficient Distributed Architecture

Abstract

Information extraction is used to automatically extract useful information from unstructured or semi-structured data. The exponential rise of unstructured, multidimensional data poses novel challenges for information extraction techniques in the era of big data. The volume and variety of big data imply the need to implement advanced techniques to extract information reliably and at scale. This paper introduces a novel approach for distributed data extraction based on the combined use of PaddleOCR and the distributed framework Apache Kafka. This work aims to present and compare four distinct methods for deploying PaddleOCR in a distributed environment, thereby integrating AI-enhanced models into a scalable and resilient framework. This article provides a roadmap for developing large-scale information extraction systems that are both high-performing and adaptable, meeting the growing demands for data processing in ever-evolving digital environments.

Research topics

  • Software System Performance and Reliability
  • Advanced Database Systems and Queries
  • Service-Oriented Architecture and Web Services

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/niss66502.2025.00027

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.