Towards Robustness in Computer-based Image Understanding

The thesis addresses current issues in the robustness of deep learning models in computer vision, a topic of growing relevance in industry and society. The Jury has assessed the feasibility, potential and originality of the proposal, as well as the freshness of the perspective with which the topic has been addressed, and the innovative solutions that have been presented to solve the problems presented. For all this, it has been considered that the findings have the potential to revolutionize the industry.

Basic Information

Diego Velazquez Dorta

Dr. Jordi Gonzalez Sabaté Dr. Josep M. Gonfaus Sitjes Dr. Pau Rodríguez López

CVC Satellogic Apple Research

Centres CERCA List
Associated Universities

CERCA Institute

Cerdanyola del Vallès, Spain

1994

Centre de Visió per Computador (CVC)

Area

DEEPTECH Area

Abstract

This thesis embarks on an exploratory journey towards robustness in deep learning, with a keen focus on the intertwined facets of generalization, explainability, and edge cases within the field of computer vision. In deep learning, robustness epitomizes the resilience and flexibility of a model, based on its ability to generalize across diverse data distributions, explain its predictions transparently, and navigate effectively through the complexities of edge cases. The challenges associated with robust generalization are multifaceted, including the performance of the model on unseen data and its defense against out-of-distribution data and adversarial attacks. To bridge this gap, we explore the potential of Embedding Propagation (EP) to improve out-of-distribution generalization, which in turn strengthens the model's robustness against adversarial attacks and improves performance in low-sample and self-supervised learning. Within the maze of deep learning models, the path to robustness often intersects with explainability. As the complexity of the model increases, so does the urgency of deciphering its decision-making processes. Recognizing this, the thesis introduces a robust framework for evaluating and comparing various counterfactual explanation methods, echoing the imperative of quality of explanation over quantity and highlighting the complexities of diversifying explanations. At the same time, the deep learning landscape is full of edge cases, anomalies in the form of small objects or rare instances in object detection tasks that challenge the norm. Faced with this, the thesis presents an extension of the DEtection TRansformer model to improve small object detection that incorporates the Feature Pyramid technique, although it faces challenges such as high computational costs. With the emergence of fundamental models in mind, the thesis unveils EarthView, the largest-scale remote sensing dataset to date, built for self-supervised learning of a robust fundamental model for remote sensing. Collectively, these studies contribute to the grand narrative of robustness in deep learning, interweaving the threads of generalization, explainability, and edge-case performance. Through these methodological advances and novel datasets, the thesis calls for continued exploration, innovation, and refinement to strengthen the bastion of robust computer vision.

The doctoral thesis that precedes this document represents a systematic effort to solve problems in the field of deep learning models, a topic of growing relevance in today's industry and society. This research is articulated through four main works. We now present a proposal for the application of the knowledge and tools developed, aimed at exploiting the commercial and industrial advantages of these discoveries and guiding the development of informed public policies. The following document explains how the results can influence and benefit key sectors. Chapter 1: Logo Detection With No Priors Research results in automatic object detection and classification, specifically through the extension of the DETR approach with feature pyramids, offer significant advances that can transform the fight against logo counterfeiting. According to the Organisation for Economic Co-operation and Development (OECD), the global trade in counterfeit goods represented up to 3.3% of global trade in 2016, amounting to $509 billion. With this large underground economy at stake, the ability to automatically detect counterfeit logos is more critical than ever for brand management and intellectual property protection. Below, we highlight some of the most promising industrial and commercial applications of our thesis results, along with quantitative data illustrating their potential impact: Commercial and Industrial Application 1. Brand Authenticity Verification: In a world where the trade in counterfeit goods is worth around $600 billion annually according to The Guardian [1], our technology is presented as a key solution for companies in sectors such as luxury goods, consumer electronics, and sportswear. The ability to verify logo authenticity can redirect these significant cash flows to legitimate businesses, improving profitability and strengthening consumer trust. 2. Counterfeit Sales Protection: With an estimated $1.7 to $4.5 trillion in counterfeit goods sold annually globally, our technology can be integrated into online marketplaces to analyze images of listed products, ensuring that logos match the authentic ones. This is vital as 52% of consumers report losing trust in a brand after purchasing a counterfeit product online [2], having a lasting negative impact on customer loyalty and sales. 1 3. Document Security Applications: Banks and insurance companies can benefit from this technology to verify the authenticity of documents and prevent fraud. This measure is of critical importance since data loss can have significant financial consequences and damage to business reputation. Document security is therefore essential to maintain customer trust and the integrity of their personal and financial data. 4. Production Quality Control: The importance of quality control in the production chain is critical, since the costs associated with errors can be substantial. Anecdotal evidence suggests that in manufacturing, the cost of poor quality often amounts to nearly 100 times the initial price of the defective part. This means that a defective $7 retainer can cost up to $700 to fix, and for a million-dollar satellite, the cost is likely to exceed $100 million [3]. Furthermore, recent studies suggest that, although manufacturers calculate their cost of quality at around 10 percent of revenue, in reality this cost is closer to 20 percent [4]. In this context, the use of this technology acquires key relevance. Within the production chain, the technology could be used to ensure that logos are applied correctly and to correct small defects in products during the manufacturing process. Chapter 2: Embedding Propagation for Manifold Smoothing With recent advances in machine learning and artificial intelligence, the need for models that generalize well in out-of-distribution (OOD) situations is more critical than ever. The results presented in the study of embedding propagation (EP) for manifold smoothing open new avenues for model generalization in commercial and industrial applications, as well as in public policy development. The application of EP can transform the way companies approach problems of image classification, object detection, and image segmentation, especially in areas where OOD data are common, such as medicine (medical imaging), security (surveillance systems), and the automotive industry (autonomous vehicles). Commercial and Industrial Application 1. Improved Robustness to Adversarial Attacks: Models that are robust to adversarial attacks are essential for critical applications such as IT security and defense. Implementing EP can improve the security of facial recognition and object detection systems in attack-sensitive environments, such as banking security or access control systems. 2. Efficiency in Data Labeling: Efficiency in data labeling is crucial in the advancement of artificial intelligence, particularly in fields such as medicine, where manual labeling can be prohibitively expensive or impractical. By using semi-supervised and self-supervised learning, powered by EP, the need for labels can be reduced, thus minimizing costs and accelerating the implementation of AI models. In the healthcare sector, this has profound implications: one report suggests that AI could improve patient outcomes by up to 40% and reduce treatment costs by up to 50% [5]. AI has already been shown to outperform human radiologists in accuracy, identifying tuberculous lesions on chest X-rays with 96% accuracy, and can diagnose breast cancer 30 times faster with 99% accuracy [5]. Furthermore, AI-driven data management optimization has led to a reduction in repeat examinations by up to 19% [3], cutting costs and reducing radiation exposure for patients 3 . These improvements not only promise substantial financial savings but also a significant advance in the quality of patient care. Chapter 3: A principled Benchmark for Visual Counterfactual Explainers Explainability in AI plays a crucial role not only in academia but also in the commercial, industrial, and political worlds, as it helps increase users' trust in AI models. When users understand how a model reaches its decisions, the likelihood that they will trust and adopt AI in various applications increases significantly. This ability to explain complex decisions has the potential to transform the way in which society views and uses AI, making it a more accessible, reliable, and valuable tool. Thus, explainability research not only serves as a foundation for future efforts to improve artificial intelligence, but also drives innovation and progress in many sectors, enabling a broader and more effective integration of AI into the fabric of daily life. Advances in the explainability of deep learning models are crucial for a wide range of commercial and industrial applications, as well as for the development of public policies. Some of the potential applications of the results obtained in this research are described below: Commercial and Industrial Applications 1. Medicine: In the field of medicine, the integration of artificial intelligence (AI) not only significantly enhances diagnosis and treatment, but also promises to reduce healthcare costs. The widespread adoption of AI could translate into savings of 5% to 10% in healthcare costs, equivalent to about 200 to 360 billion dollars annually [3]. In addition, the implementation of this technology leads to a reduction in the time spent on treatments, starting with a saving of 21.67 hours per day per hospital in the first year, and reaching a peak of 122.83 hours per day per hospital in the tenth year. This increase in efficiency suggests a considerable cost reduction over time [4]. However, for this acceptance to occur, it is essential that AI systems provide clear and understandable explanations, which will foster the trust of professionals and patients. Explanatory models, which improve the communication of clinical decisions, must also comply with strict regulatory requirements that demand transparency in medical technologies. This combination of economic efficiency and clarity of reasoning is key to the successful integration of AI into public health. 2. Finance: In the financial sector, transparency and trust are essential for decision-making in loans and investments. Detailed explanations generated by AI models can not only improve customer relationships, but also ensure regulatory compliance, a fundamental aspect in a highly regulated sector. This clarification in AI-assisted decisions can also facilitate internal risk management and the justification of financial strategies. 3. Automotive: In the automotive industry, especially with the advancement of autonomous vehicles, the explainability of decisions made by artificial intelligence (AI) models is becoming a critical factor in increasing user trust and accelerating regulatory approval. A clear understanding of how and why an autonomous vehicle makes specific decisions in real time is not only essential for ensuring safety, but can also have a significant economic impact: a 1% reduction in road safety incidents in the US could save more than $8 billion annually [5]. This underscores the potential of safe, AI-powered autonomous vehicles to transform not only mobility, but also the economy. This holistic view of the technology emphasizes the importance of public acceptance and trust in this new era of mobility being redefined by AI. Public Policy Development 1. Transparency and Bias: As governments begin to adopt AI for decision-making, it is essential that these models are transparent. This can help to prevent biased decisions and ensure that citizens' rights are respected. Chapter 4: EarthView: A Large Scale Remote Sensing Dataset Significant progress in artificial intelligence modeling and the accumulation of large datasets have unlocked new possibilities in the field of remote sensing. In particular, EarthMAE emerges as a fundamental model in this sphere, distinguishing itself by its outstanding ability to execute novel tasks, showing a remarkable performance even in zero-shot scenarios, where the model is applied to tasks that it has not seen during training. This indicates its versatility and its powerful potential for extraction of knowledge in situations not previously anticipated. At the same time, EarthView presents itself as a large-scale dataset (20 trillion pixels), unprecedented in its breadth and depth, enabling an extensive application in any terrestrial monitoring task. The combination of a model like EarthMAE, with its exceptional generalization capabilities, and a dataset like EarthView, rich in variety and volume, creates a perfect symbiosis for innovation in multiple fields. This opens the door to the exploration and application of these resources in a wide range of Earth monitoring tasks, including those detailed below: Commercial and Industrial: 1. Precision Agriculture: Companies in the agricultural sector can benefit from a 4% increase in crop production and a 7% increase in fertilizer placement efficiency thanks to precision agriculture technologies [6]. These advances allow not only to optimize crop yield and early detection of pests and diseases, but also to manage water resources more efficiently. The result is a significant improvement in the sustainability and efficiency of food production, with the addition of a more than 10% reduction in carbon footprint [7], which can have an impact on direct and indirect economic savings, including the possible obtaining of carbon credits. 2. Forest Management and Environmental Monitoring: The application of active forest management can have notable benefits in carbon capture [8], surpassing unmanaged forests, and provides the additional advantage of generating low-emission energy from forest and sawmill residues. In regions such as California, where more than 1,000 million dollars are spent annually on fighting wildfires [8] , a proactive implementation of forest management practices could imply savings of hundreds of millions, reducing the fuel load and the severity of fires. 3. Urban Planning and Infrastructure: In the urban context, trees and non-motorized transport infrastructure play a key role. For example, in Tshwane, South Africa, planting 1,000 additional trees has been estimated to save $3.82 million in annual cooling energy [9], and in Coimbatore, India, an investment in non-motorized transport infrastructure is expected to generate net benefits of up to $510 million over 23 years [10]. These data underscore the importance of efficient infrastructure planning and the creation of optimized spaces that can result in more sustainable and green cities, with a return of $31 in benefits for every dollar invested, thus fostering environmentally conscious and respectful urban expansion. Public Policy Development: 1. Climate Change Mitigation: Governments could use insights to formulate more informed policies on climate change, such as improving carbon emissions management or planning reforestation initiatives. 2. Natural Resource Management: It could improve the management of water, forests and biodiversity, providing tools for data-based decision-making for the conservation and sustainable management of natural resources.

Robustness; Deep Learning; Generalization; Explainability; Edge Cases; Computer Vision; Resilience; Flexibility; Data Distributions; Transparent Predictions; Unseen Data; Out-of-Distribution Data; Adversarial Attacks; Embedding Propagation (EP); Low-Sample Learning; Self-Supervised Learning; Counterfactual Explanation Methods; Quality of Explanation; Object Detection; Small Object Detection; Rare Instances; DEtection TRansformer (DETR); Feature Pyramid Technique; High Computational Costs; Fundamental Models; EarthView; Remote Sensing Dataset; Self-Supervised Learning; Robust Computer Vision.