Domain adaptation between virtual and real worlds applied to pedestrian detection

Simplification of the product improvement process for road safety and other industrial applications.

High proximity to the market.

Basic Information

David Vázquez Bermúdez

Lopez Peña, Antonio Manuel Ponsa, Daniel

ADAS - Advanced Driver Assistance Systems (CVC)

Centres CERCA List
Associated Universities

https://portalrecerca.csuc.cat/107343295

CERCA Institute

Cerdanyola del Vallès, Spain

1994

Centre de Visió per Computador (CVC)

Area

DEEPTECH Area

Abstract

Pedestrian detection is key to many applications such as driver assistance, video surveillance or multimedia. The best detectors are based on classifiers based on models trained appearance with annotated examples. However, the annotation process is a chore intensive and subjective when carried out by people. Therefore, it is worth minimizing the human intervention in said task through the use of computational tools such as virtual worlds because with them we can obtain varied and precise annotations quickly. However, the use of this type of data generates the following question: Is it possible that a model of appearance trained in a virtual world can work satisfactorily in the real world? To answer this question, we have carried out different experiments that suggest that Classifiers trained in the virtual world can offer good results when applied real world environments. However, it was also found that in some cases these classifiers can be affected by the problem known as the change in the nature of the data, just as it happens with the classifiers trained in the real world. Consequently, we have designed a domain adaptation system, V-AYLA, in which we have tested different techniques to collect a few examples from the real world and combine them with a large number of examples from the virtual world to train a pedestrian detector adapted V-AYLA offers the same detection accuracy as a trained detector manual annotations and tested with real images from the same domain. ideally, we would like our system to adapt automatically without the need for human intervention. Therefore, as a demonstration, we propose to use techniques of unsupervised adaptation that allows to completely eliminate human intervention of the adaptation process. As far as we know, this is the first work he shows that it is possible to develop an object detector in the virtual world and adapt it to the world real Finally, we propose a different strategy to avoid the change problem in the nature of the data that consists of collecting examples in the real world and retrain only with them but doing it in such a way that they don't have to write down pedestrians in the real world. The result of this classifier is equivalent to another trained with annotations obtained manually. The results presented in this thesis is not limited to adapting a virtual pedestrian detector to the real world, but that goes further, showing a new methodology that would allow a system to adapt to any new situation and that lays the foundations for future research in this still unexplored field.

This thesis addresses different important problems for the industry: detection of pedestrians, the generation of synthetic data and domain adaptation. The detection of people/pedestrians is of great importance in the automotive industry both in the driving assistance systems to prevent accidents and in the autonomous navigation It also has many other applications in such diverse industries such as security, audiovisual, entertainment, smartphones, etc. us We propose a robust person detection method that we have successfully applied in industrial projects with two important companies in the automotive sector and one with a smartphone manufacturer. Any application based on computer vision requires data from training to learn a model. These data must contain a large amount of examples that reflect the different variations that can present an object and they must be correctly noted. Obtaining all this information is not an easy task It tends to be very tedious and involves a great economic effort when developing one industrial application. We propose the use of virtual images from a video game to solve this problem. Most applications based on computer vision face the problem of generalization since the conditions in which the data were acquired of training usually vary with respect to the conditions of the different environments of application We propose the application of domain adaptation techniques to solve this problem. Companies like Xerox have shown interest in our system

Pedestrian Detection; Driver Assistance; Video Surveillance; Multimedia; Detectors; Classifiers; Models; Trained Appearance; Annotated Examples; Annotation Process; Chore Intensive; Subjective; Human Intervention; Computational Tools; Virtual Worlds; Varied Annotations; Precise Annotations; Quick Annotation; Data Generation; Appearance Model; Real World; Experiments; Good Results; Real World Environments; Classifiers; Change in the Nature of the Data; Domain Adaptation System; V-AYLA; Techniques; Real World Examples; Virtual World Examples; Pedestrian Detector; Detection Accuracy; Manual Annotations; Real Images; Automatic Adaptation; Unsupervised Techniques

Content Blocks

Awarding Ceremony