Theses supervised by Dr. Öğr. Üyesi Funda Yıldırım
10 theses · Yeditepe University
The effect of parametric physical attributes of 2d shapes on shape recall and on classifiying shapes as natural or unnatural
Previous research in the visual perception domain has shown the interaction between some attributes of shapes and how humans perceive them differently when those attributes are changing. However we felt there are not enough resources in the literature regarding the geometrical parameters of 2D shapes. Also about the naturalness classification in visual perception, there are studies focusing on the natural scenes but again, the knowledge of the naturalness perception according to the 2D geometrical shape parameters is pretty limited. In this study we aimed at combining these two subdomains and providing quantitative findings about how the geometrical parameters like edge count, concavity, naturalness and uniformity are affecting the memory performance and naturalness classification for a shape. We designed a behavioural experiment to test different shapes with different values of those parameters and presented these shapes to participants in a recall task and a naturalness survey. Then we ran some correlation analysis between those values and memory performance and naturalness rating. We found out that edge count and concavity significantly affects the memory performance for a shape. Another finding was that our judgement of a shape being natural or artificial is affected by the edge count, concavity level and roundness level of a shape. Lastly, we observed that the memory performance is worse for the shapes that are rated as natural.
The effect of contour and configuration on visual crowding
Visual crowding refers to poor target identification performance due to the presence of flankers. Even though crowding studies mostly presented for low level stimulus settings, it is observed in high level settings as well such as faces and configuration. This study is aimed to investigate the effect of contour type (sharp-edged vs curve-edged) and configuration type (random vs smooth) with two experiments in crowded scenes. We hypothesized that the sharp-edged contours should be reacted faster for and more accurately compared to the curve-edged ones due to assumption of hierarchical and pooling models. According to both models, objects are recognized in a hierarchical way such that low-level and high-level features are orderly processed in the cortex. Sharp-edged contours might have an identification advantage due to comprising simple lines and intersections whereas circle does not have obvious components as the triangle has. We also expected that random configuration trials to have better outcome rates compared to regular and irregular ones because grouping can make detection of the target harder due to the similarity between target and flankers. We performed two experiments where we controlled configuration and contour respectively. We compared reaction times (ms) and accuracy rates (percentage) as outcome variables. Results of the first experiment have confirmed the hypothesis that the advantage of sharp-edged contour was obtained from both the accuracy and reaction time analysis. Also, random configurations were responded more accurately, as we had expected. However, main effect of configuration was not reflected on the reaction times. The second experiment was motivated to find the impact of high-level stimulus features by adding an incongruent contour type to smooth and regular configurations. Our results indicated the advantage of sharp-edged contours and random configurations in crowded scenes for identification performance.
The effect of harmonic and inharmonic sound on roughness perception in the tongue and fingertip
Multisensory integration refers to the integration of multiple senses by the nervous system. Auditory and tactile features are closely related senses as can be understood from the fact that adjectives such as soft, rough, and warm are used commonly for auditory and tactile features. Previous studies show that different characteristics of auditory cues may cause perceiving tactile cues as smoother or rougher. In this thesis, the effect of harmonic and inharmonic sounds on tactile roughness perception was studied by tongue and fingertip while they are presented simultaneously. We expected the participants to perceive surfaces rougher while they listen to inharmonic sounds due to auditory roughness. We presented simultaneous and sequential harmonic and inharmonic sounds with three sandpapers with different roughness levels (P100, P120, P150 grit numbers) to the participants. Results suggest that participants perceive sandpaper with the P120 grit number rougher while they listen to simultaneous inharmonic sounds than simultaneous harmonic sounds by their fingertip in two different experiments. However, any effect of harmonicity on the sandpapers with P100 and P150 grit numbers was not observed for fingertip. In addition to this, there is no effect of harmonic and inharmonic sounds on roughness perception for tongue. We suggest that auditory roughness may enhance tactile roughness perception for surfaces with particular roughness levels for fingertip, possibly when there is an ambiguity in roughness estimation.
Investigating the role of position and orientation in shape perception using visual crowding and contour integration
The deterioration or reduction of the perception of the target due to the nearby objects is called the visual crowding effect. Crowding can occur as a result of similarity and proximity. These two conditions also result in contour perception with high congruency between the elements. Contour perception is the unitary perception of congruent individual items. Although many physical aspects can result in crowding and contour perception, congruencies of orientation and position are the two primary indicators. In this experiment, four positional configurations, namely left-circle, right-circle, left-diagonal and right-diagonal, of the target and flankers were used. According to these positional configurations, congruent and incongruent orientations were created. We hypothesized that crowding and contour should have an exhibitory effect on each other. In the previous studies, target and contour objects were separated from each other. In this study, the target was also a part of the contour pattern. This allowed us to observe the interaction of crowding and contour in the same arrangement. The task was a simple orientation discrimination task between two targets. Another hypothesis was that low-level shapes' contour perception should occur even with three objects in the scene. The accuracy rates (percentage) and reaction times (ms) were compared as outcome variables. The results indicated that contour and crowding are two concepts that cannot co-exist based on the accuracy results. Also, the analysis of the accuracy rates confirmed that the three objects were enough to create simple linear shapes but not circular shapes.
Examining the role of eye movements in facial expression and emotional intensity recognition via machine learning
Recognition and classification of emotional face expressions is one of the most advanced abilities of humans. Many studies have investigated recognition processes and scanning strategies for different facial expressions but there is still no consensus on a particular scanning pattern for emotion types. Most of the previous studies also did not include the level of emotional intensity of facial expressions. The aim of this study is to investigate which eye movement parameters are important for emotional expression recognition and emotional intensity recognition processes. Eye movements provide important clues about face scanning strategies. Therefore, we examined the eye movements of 60 participants and fed machine learning models with eye tracking data. Neutral, surprised, sad, happy, and fearful facial expressions were presented to the participants with 0%, 25%, %50, %75, and %100 emotional intensity. Half of the participants were asked to classify the emotion type and half of them were asked to rate the intensity level of the emotion. Total number of fixations, pupil size and number of fixations on areas of interest (AOIs: eyes, nose and mouth) were collected for eye tracking analysis. Results showed that the fixation number on facial expressions decreased as the intensity of the expression increased. It was also found that participants fixated most to the eyes and least to the mouth regardless of the task. There was no difference between fixation numbers to eyes and nose during intensity recognition tasks for all emotion types. However, the fixation number towards the eye, nose and mouth differed between emotion types during the emotion recognition task. Higher emotional intensity resulted in less fixations in the mouth area. Analysis of pupillary response showed that fearful faces resulted in larger pupil size compared to all other expressions during the emotion recognition task but this difference was not evident during the intensity recognition task. Data collected from the eye tracking device has also been implemented in Random Forest (RF), Decision Tree and k-Nearest Neighbors (k-NN) models. All models achieved an accuracy rate at above chance level for both emotion classification and intensity classification. The highest accuracy rate was achieved with the RF method. The model was able to classify emotion types with an average accuracy of 85.8% for emotion recognition task and with an average accuracy of 89.2% for intensity recognition task when implemented with fixation point of gaze, fixation duration and pupil size. Emotion and intensity classification remained similar for the intensity recognition task with the same parameters. Effect of task, emotion type and intensity level were observed for pupil size and fixation numbers in the eye tracking analysis. However, the task did not affect the classification performance of the model. Also, intensity level of expressions or emotion type of an intensity level did not affect machine's performance as well. We interpreted these results as scanning strategies implemented by viewers while reading emotional expressions with different intensities. Results of the current study further discussed possible application areas.
Effects of stimulus type on eye movement measures during mindless reading and mindless listening
"Mind wandering" refers to a state when attention is directed to internal thoughts unrelated to the sensations and perceptions of the immediate physical environment. Mind wandering occurs not only in passive situations but also during tasks that require attentional effort such as reading and listening. Despite its ubiquitous nature, the number of studies which examine mind wandering has increased only in the last few decades, particularly due to its subjective nature and spontaneous occurrence. Also, the majority of those studies focused on mind wandering during reading and investigated the differences between "mindless reading" and normal reading. To the best of our knowledge, there are no studies which examine mind wandering in different modalities, the effect of stimulus types on the frequency of mind wandering or contrast the behaviour of eyes under the effect of stimulus types during mind wandering. In this study, we analysed eye movements and pupil size in three conditions, that is, during only reading, only listening and simultaneous reading and listening, in three intervals preceding and following self-reported mind wandering episodes. The results of our study revealed that mind wandering frequency is not affected by stimulus type or cognitive load. The dominance of inner thoughts affects the frequency of mind wandering more than modalities, cognitive load and the demands of the here and now. Analysis of eye-tracking data showed that there are no differences in eye movements and pupil size between mindless and attentive states. However, while fixation duration and pupil size did not vary for modalities and intervals, the number of fixations were significantly higher during only reading than during listening. Thus, our study shows that the absence of any significant differences in fixation count and fixation duration between attentional states shows that a low-level processing of visual-lexical cues continues during mind wandering. These results indicate a sensitivity to visual-lexical stimulus even in absence of conscious effort and an attentional decoupling mechanism in favour of visual-lexical stimulus. Keywords: Attentional decoupling, eye tracking, meta-awareness, mind wandering, mindless listening, mindless reading
Auditory inattentional deafness investigated with eye tracking
This thesis will examine auditory selective attention and inattentional deafness with an eye-tracking approach. The current experiment aims to take the first step and expand the understanding of the effects of multimodal changes on inattentional deafness. The main aim of this study is whether the type of environmental sounds affects the recognition of hidden sounds and whether the participants counting the lightning strike (thunder) and the sound of the wave hitting the shore will have an impact on the inattentional deafness or not. Our research questions are; (1) the participants who hear the negative environmental sound (Thunder) and the positive environmental sound (Wave) will be less likely to detect the hidden sound than the participants who hear the hidden sounds in silent conditions, but also we hypothesis that the environmental negative sound will decrease the awareness of the hidden sounds depending on its effects on participants internal mood, (2) participants pupil size will show the difference when hidden sounds are given in the experiment. Results of the experiment showed that the sound type had an impact on participants' awareness of the hidden sounds. Besides, when they hear the hidden sounds during the experiment, their pupil dilation and pupil constriction occurred. Our results indicated that hidden sound awareness affects the pupil size of the participants in general and the type of environmental sound had an impact on hidden sound awareness. Keywords: Environmental sounds, inattentional deafness, selective auditory attention,
Motion perception in multiclustered environments
The ability to recognise objects in the environment is a fundamental process for all living beings. The features of the objects (size, color, location, etc) in addition to their integration with the environment play a significant role in identifying the object. Despite the extensive evidence in psychology and neuropsychology, the exact contribution for various features to recognition is still inconclusive. Although it is an easy task for the brain to identify an object, it remains difficult to identify an object in a crowded scene – a perceptual phenomenon where having similar flankers around a target object decreases the visual acuity of the target. The more the objects are similar to each other, the more difficult it is to identify the target object. In this thesis, we focused on studying the effect of global and local features in a crowded scene of multiple clusters on detecting the target object. To achieve this, similar stimulus metrics (e.g. density, distance, size) are calculated globally for the entire scene and locally for the target object and/or cluster. We found a global effect of the density metrics (e.g. the number of clusters) on performance. We also found a significant effect on global and local aspects of eccentricity and size metrics. Keywords: Motion Perception , Ensemble Perception, Attention
Recognition of sequential harmonic and inharmonic sound sequences investigated by behavioral and electrophysiological studies
Consonance rating when listening to the two consecutive or simultaneously presented sounds is associated with the ratio of their fundamental frequencies. The simpler ratio is perceived as more consonant. In this thesis, two melodies that consist of the same 10 tones were generated. Melodies had the same contour but differed with the order of tones so formed different harmonic relations between consecutive tones. Non musician participants were asked to choose which melody (harmonic or inharmonic tone sequences) was more consonant and the rating for each melody was not significantly different than each other. However, it was easier for them to notice the same amount of frequency change that occurs in one note of the harmonic sound sequence. Results suggest that better performance of the recognition task for harmonic sequence is related to better encoding and building a higher level expectation due to harmonic relations of sounds. An EEG experiment was conducted to test memory load hypothesis. The sustained anterior negativity (SAN) and power of alpha/beta oscillations were compared as a measure of working memory load. Amplitude of SAN did not reveal a difference in working memory load. However, beta and low gamma power was lower for harmonic sequences in the frontal cortex at the end of the stimulus presentation which suggests a more stable memory trace for harmonic sequences. Nonmatch stimulus elicited a P3b signal with higher amplitude for harmonic condition. This result suggests that a higher level expectation occurs for harmonic condition as all the other parameters that affect P3 signal were well controlled in the experiment.
E-commerce product matching with deep learning
As the E-Commerce market grows, more products gets listed everyday. This growth in number of choices generate novel problems that are needed to be solved by businesses like price comparison sites and e-commerce competition analytics platforms. One such problem is matching same products despite the differences in representation. These differences can occur as differently written titles or usage of synonyms of the same product specifications. For matching products despite these different representations, we present two solutions, a metric-learning based solution for search and retrieval of the products and a siamese deep neural network for comparing product representations. Both of these models only needs product titles and are specialized for Turkish language.