MARATTO

article

Unlocking Arabic Script: OCR Technology for Efficient Text Extraction

Abstract

Optical Character Recognition (OCR) has evolved as a key technique for converting handwritten text and digits into editable and searchable digital representations. In this study, we investigate OCR for Arabic language documents, with an emphasis on handwritten letters and numerals. Despite developments in OCR, the complexity of Arabic script provide hurdles for reliable recognition. Our project intends to investigate post-processing approaches for improving the quality and usefulness of OCR results for handwritten Arabic content. Using sophisticated neural network designs, notably Convolutional Neural Networks (CNNs), we examine picture preprocessing, language-specific corrections, and layout analysis to improve OCR performance. Through empirical assessment, our model accomplished amazing accuracies of 98.5 percent for Arabic digits and 97.8 percent for Arabic letters, illustrating the viability of our comprehensive approach to digitizing manually written Arabic records and contributing to the progression of OCR innovation for multilingual applications.

Research topics

  • Handwritten Text Recognition Techniques
  • Natural Language Processing Techniques
  • Mathematics, Computing, and Information Processing

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/imsa61967.2024.10652687

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.