
Document Image Enhancement and Recognition in Low Resource Scenarios: Application to Ciphers and Handwritten Text
Basic Information
Mohamed Ali Souibgui
2023
Alicia Fornés Yousri Kessentini
CVC Digital Research Center of Sfax (CRNS).
Prize
Male
CVC
Universitat Autònoma de Barcelona (UAB)
CERCA Institute

Cerdanyola del Vallès, Spain
1994
Centre de Visió per Computador (CVC)
Area
Automation
Novel AI
Industry
Information & communication technology
Machine Learning & Artificial Intelligence
Abstract
This thesis introduces innovative contributions aimed at advancing the enhancement and recognition of handwritten document images, in particular those that include rare alphabets such as encrypted documents or highly degraded documents. Beyond the scope of technical advances, our research explores the profound impact on the optical character recognition (OCR) industry, recognizing the tangible applications of these innovations. In the first part of the thesis, we focus on solving the challenge of paper degradation, an obstacle for current OCR systems. To solve this, we present end-to-end models for document image enhancement (DIE) using deep learning techniques. This involves deploying generative adversarial networks (cGANs) for tasks such as document cleaning, binarization, blurring, and removing watermarks on paper. In particular, we integrate a text recognition module into the cGAN model, improving the quality of images and making them clearer and more readable. In addition, the introduction of a new encoder-decoder architecture based on vision transformers provides a comprehensive end-to-end solution for enhancing document images, both printed and handwritten. This advancement makes it easier for OCR systems to accurately recognize and convert text from images into editable and searchable digital formats, thus impacting several industrial fields. The second part of the thesis focuses on text recognition in resource-poor scenarios, precisely one of the challenges industries face when there is little labeled data to train deep learning models. Specifically, we propose methods for recognizing rare alphabets with few data, including a system based on few-shot object detection (fOL) and a progressive learning strategy to minimize human annotation efforts while maintaining model performance. In addition, we introduce a data generation technique based on Bayesian Program Learning (BPL) to compensate for the sparse labeled data. Finally, we present a Text Degradation Invariant Autoencoder (Text-DIAE), a self-supervised model that simultaneously addresses text recognition and document image enhancement, requiring substantially fewer data samples to converge, a critical efficiency factor for OCR and machine learning industry applications.
In summary, OCR has been a key component of computer vision (CV), driven by deep learning. The thesis concentrated on enhancing OCR through image improvement and data-efficient training in resource-constrained scenarios. Advancements in image enhancement, employing cGANs and vision transformers, enhanced OCR accuracy and efficiency in diverse sectors. Innovations like few-shot learning, data generation, and self-supervised learning addressed limited training data challenges, extending OCR to less common text images. These developments continue to boost productivity and accessibility across industries, aligning with the evolving digital landscape.
Handwritten Document Images; Enhancement; Recognition; Rare Alphabets; Encrypted Documents; Highly Degraded Documents; Optical Character Recognition (OCR) Industry; Technical Advances; Tangible Applications; Paper Degradation; OCR Systems; End-to-End Models; Document Image Enhancement (DIE); Deep Learning Techniques; Generative Adversarial Networks (cGANs); Document Cleaning; Binarization; Blurring; Watermark Removal; Text Recognition Module; Image Quality Improvement; Readability; Encoder-Decoder Architecture; Vision Transformers; Printed Documents; Handwritten Documents; Industrial Fields; Text Recognition; Resource-Poor Scenarios; Labeled Data Scarcity; Few-Shot Object Detection (fOL); Progressive Learning Strategy; Human Annotation Efforts Minimization; Model Performance Maintenance; Data Generation Technique; Bayesian Program Learning (BPL); Sparse Labeled Data Compensation; Text Degradation Invariant Autoencoder (Text-DIAE); Self-Supervised Model; Simultaneous Text Recognition; Simultaneous Document Image Enhancement; Reduced Data Samples; Critical Efficiency Factor; OCR Industry Applications; Machine Learning Industry Applications.