Master'sOpen Access

Enhancing scene sketch understanding through a dual-network: Visio-temporal segmentation and context-aware sketch recognition

2024
0 views
0 downloads
Advisor: Prof. Dr. Tevfik Metin Sezgin

Abstract (EN)

Understanding scene sketches involves segmenting and categorizing individual objects within the sketch. Semantic segmentation in scene sketches is crucial for distinguishing distinct sketches, but current methods often treat sketches as bitmap images, which can result in a loss of stroke order information. However, people tend to draw objects sequentially, so leveraging this temporal order could enhance segmentation performance. Moreover, traditional methods typically focus on class-level segmentation, failing to differentiate between instances within the same category. Another important aspect of scene sketch understanding is classifying individual sketch objects within a scene. These objects often lack the detail needed for standalone recognition, making their identification challenging without contextual information. For instance, a sketch that is fluffy and circular could be interpreted as a "bush" or a "cloud," depending on its position and size within the scene. Bushes are typically drawn on the ground, while clouds are usually sketched in the sky. Despite their similar appearances, their interpretation can vary based on their context. To enhance recognition accuracy, information about relative position and size is often used in computer vision tasks like object recognition and image classification. However, many current sketch recognition methods treat sketches in isolation, overlooking the contextual information present in the scene. To address these issues, I propose a dual-network approach comprising two novel networks for separate tasks: scene sketch segmentation and scene sketch recognition. The first network, the Class-Agnostic Visio-Temporal Network (CAVT), detects individual objects in a scene sketch using a class-agnostic object detector and groups strokes with its post-processing module. This network can distinguish object instances at the stroke level, independent of their categories. The second network, Context-Aware Graph Attention Transformer Network (CGAT-Net), processes individual sketch objects within the scene and leverages inter-object relationships to find their appropriate categories. This work is the first to apply a context-based sketch recognition approach by leveraging a novel Transformer-based Graph Attention Network within scene sketches. Additionally, the literature lacks free-hand scene sketch datasets with both instance and stroke-level class annotations. To fill this gap, I collected the largest Free-hand Instance- and Stroke-level Scene Sketch dataset (FrISS) that contains 1,000 scene sketches and covers 403 different object classes with dense annotations. Extensive experiments on FrISS and other scene sketch datasets demonstrate that the dual-network approach, combining CAVT and CGAT-Net, as well as each network individually, outperforms existing methods in their respective domains.

Author

Dr. Aleyna Kütük

How to Cite

Aleyna Kütük (Master Thesis). Enhancing scene sketch understanding through a dual-network: Visio-temporal segmentation and context-aware sketch recognition, 2024, Koç University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University