Millora i reconeixement d'imatges de documents en escenaris de baixos recursos: aplicació a xifratges i text manuscrit

La tesi presenta contribucions innovadores destinades a avançar en la millora i el reconeixement d'imatges de documents manuscrits, especialment aquells escrits en alfabets rars o poc comuns per als quals hi ha poques dades etiquetades disponibles, mitjançant tècniques d'aprenentatge profund, alhora que explora l'impacte en la indústria del reconeixement òptic de caràcters. El jurat ha valorat l'enfocament innovador de la tesi per abordar el repte que plantegen les dades d'entrenament limitades, així com el potencial de la tecnologia per a la seva viabilitat i aplicació en el món real en casos on abans era molt difícil aplicar el reconeixement òptic de caràcters.

Informació básica

Mohamed Ali Souibgui

Alicia Fornés Yousri Kessentini

direcció d'Alicia Fornés, investigadora sènior del Centre de Visió per Computador – CVC; i Yousri Kessentini, Centre de Recerca Digital de Sfax (CRNS).

Llista Centres CERCA
Universitats associades

https://portalrecerca.csuc.cat/107417833

Institut CERCA

Cerdanyola del Vallès, Spain

1994

Centre de Visió per Computador (CVC)

Àrea

Àrea DEEPTECH

Resum

This thesis introduces innovative contributions aimed at advancing the enhancement and recognition of handwritten document images, in particular those that include rare alphabets such as encrypted documents or highly degraded documents. Beyond the scope of technical advances, our research explores the profound impact on the optical character recognition (OCR) industry, recognizing the tangible applications of these innovations. In the first part of the thesis, we focus on solving the challenge of paper degradation, an obstacle for current OCR systems. To solve this, we present end-to-end models for document image enhancement (DIE) using deep learning techniques. This involves deploying generative adversarial networks (cGANs) for tasks such as document cleaning, binarization, blurring, and removing watermarks on paper. In particular, we integrate a text recognition module into the cGAN model, improving the quality of images and making them clearer and more readable. In addition, the introduction of a new encoder-decoder architecture based on vision transformers provides a comprehensive end-to-end solution for enhancing document images, both printed and handwritten. This advancement makes it easier for OCR systems to accurately recognize and convert text from images into editable and searchable digital formats, thus impacting several industrial fields. The second part of the thesis focuses on text recognition in resource-poor scenarios, precisely one of the challenges industries face when there is little labeled data to train deep learning models. Specifically, we propose methods for recognizing rare alphabets with few data, including a system based on few-shot object detection (fOL) and a progressive learning strategy to minimize human annotation efforts while maintaining model performance. In addition, we introduce a data generation technique based on Bayesian Program Learning (BPL) to compensate for the sparse labeled data. Finally, we present a Text Degradation Invariant Autoencoder (Text-DIAE), a self-supervised model that simultaneously addresses text recognition and document image enhancement, requiring substantially fewer data samples to converge, a critical efficiency factor for OCR and machine learning industry applications.

In summary, OCR has been a key component of computer vision (CV), driven by deep learning. The thesis concentrated on enhancing OCR through image improvement and data-efficient training in resource-constrained scenarios. Advancements in image enhancement, employing cGANs and vision transformers, enhanced OCR accuracy and efficiency in diverse sectors. Innovations like few-shot learning, data generation, and self-supervised learning addressed limited training data challenges, extending OCR to less common text images. These developments continue to boost productivity and accessibility across industries, aligning with the evolving digital landscape.

Imatges de documents manuscrits; Millora; Reconeixement; Alfabets rars; Documents xifrats; Documents altament degradats; Indústria del reconeixement òptic de caràcters (OCR); Avenços tècnics; Aplicacions tangibles; Degradació del paper; Sistemes OCR; Models d'extrem a extrem; Millora d'imatges de documents (DIE); Tècniques d'aprenentatge profund; Xarxes generatives antagògiques (cGAN); Neteja de documents; Binarització; Desenfoque; Eliminació de marques d'aigua; Mòdul de reconeixement de text; Millora de la qualitat de la imatge; Llegibilitat; Arquitectura codificador-descodificador; Transformadors de visió; Documents impresos; Documents manuscrits; Camps industrials; Reconeixement de text; Escenaris amb pocs recursos; Escasesa de dades etiquetades; Detecció d'objectes de pocs trets (fOL); Estratègia d'aprenentatge progressiu; Minimització dels esforços d'anotació humana; Manteniment del rendiment del model; Tècnica de generació de dades; Aprenentatge de programes bayesià (BPL); Compensació de dades etiquetades disperses; Autocodificador invariant a la degradació del text (Text-DIAE); Model autosupervisat; Reconeixement de text simultani; Millora simultània d'imatges de documents; Mostres de dades reduïdes; Factor d'eficiència crítica; Aplicacions industrials d'OCR; Aplicacions industrials d'aprenentatge automàtic.