High resolution facial image acquisition with reinforcement learning
2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Burhan Baraklı
Abstract (EN)
This study proposes an innovative approach that combines deep reinforcement learning (DRL) and a residual dense blocks in residuals (RRDB) based method to convert low-resolution face images into high-resolution ones. Nowadays, facial super resolution (FSR) is of great importance in many critical areas such as biometric security, face recognition systems, forensics, healthcare, digital media and entertainment industry. The enhancement of low-quality facial images is becoming a critical requirement, especially in analyzing data from a variety of sources such as surveillance cameras, old photographs, and low-resolution digital images. Existing super-resolution techniques are largely based on convolutional neural networks (CNN), generative adversarial networks (GAN) and deep learning-based methods. However, most of the methods do not take into account the loss of detail in certain regions of the face images and process the whole image equally. Thus, detail loss may occur in certain critical regions of the face images and the overall visual integrity may be compromised. In the proposed work, we develop a super-resolution model supported by a reinforcement learning-based attention mechanism that identifies important regions of facial images and iteratively enhances the selected parts. The proposed model makes the super-resolution problem a dynamic process by using a mechanism that learns which regions of the face images need to be enhanced. Traditional methods approach the whole face image in the same way and try to improve the detail. The proposed model has a decision-making process that learns that different regions in the face image should be enhanced at different levels of detail. The attention mechanism of the model iteratively selects certain regions in the face image and reconstructs them with a super-resolution process at each iteration. The process continues by enhancing each new face region and integrating it with the regions enhanced in the previous iterations to maintain the global structure of the face image. One of the key components of the proposed model is a reinforcement learning-based strategic part selection mechanism. By identifying critical details in the face image, the mechanism enables the model to understand which regions need further enhancement. While traditional super-resolution methods process images as a whole, the proposed model uses the attention mechanism to focus on a specific facial region and perform enhancement at each iteration. At each step, the model achieves the ability to analyze the missing details in the face image. The model identifies the regions that need to be improved and integrates them back into the image by detailing these regions with an RRDB-based super-resolution network. The process of the model here continues with the selection and enhancement of a new face region in each iteration. The goal is to obtain an improved high-resolution face image. One of the major innovations of the model is the stochastic action selection mechanism. Traditional super-resolution methods perform image enhancement by acting according to deterministic (predetermined) rules. The proposed model uses a stochastic (probabilistic) process to decide which facial region to enhance. The stochastic process allows the model to learn the best enhancement strategy for each image. Thanks to the stochastic action selection, the model is able to identify the most critical details in different facial structures and create an enhancement strategy specific to each image. Thus, instead of a fixed enhancement algorithm, the model provides a super-resolution method that can be dynamically shaped according to the different requirements of each face image. The super-resolution phase of the model is realized using a residual dense blocks within residuals (RRDB) based network. The RRDB structure efficiently recovers missing details in low-resolution face images through dense connections and residual learning mechanisms. In each iteration, the model increases the details of the selected face region to high resolution. It then integrates the selected region into the low-resolution face image. Thus, the overall structure of the face is preserved and the most important details are enhanced. The performance of the model is evaluated by testing it on the widely used face image datasets LFW, CelebA, BioID and PubFig. The results show that the model outperforms existing super-resolution methods both visually and numerically. In particular, the analysis using metrics such as PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure) shows that the model is very successful in preserving the details in the face images and improving the global structure. PSNR is a metric that measures the level of error between two images and numerically expresses how small the difference between the low-resolution image and the high-resolution predicted image is. Higher PSNR values indicate a lower error rate and hence the image is closer to high resolution. However, the PSNR metric alone may not be sufficient to fully express the quality of an image. Therefore, in applications that require a high level of visual detail, such as super resolution, it is critical to support it with the SSIM metric. SSIM is a metric that evaluates image quality by considering not only the direct difference between pixels but also the structural similarities, brightness, contrast and edge details in the image. Since human visual perception is more sensitive to the preservation of structural information rather than absolute pixel differences, the SSIM metric is considered one of the most critical quality measures in super-resolution evaluations. Especially in face images, preserving the details of critical regions such as eyes, nose and mouth is one of the most important factors determining the super-resolution success. The model proposed in this study shows a superior performance compared to the existing super-resolution methods by having a low error rate in terms of PSNR and improving the details by preserving the spherical structure of the face with a high SSIM value. In particular, problems such as blurred edges, artificial detail addition and distortion of the natural structure of the face, which are frequently encountered in super-resolution face images produced by conventional methods, are minimized thanks to the attention mechanism and iterative super-resolution approach of the proposed model. SSIM evaluations prove that the model preserves the integrity of the face, enhances the details in the best way and produces results that are more suitable for human perception. The proposed work improves the applicability of reinforcement learning-based super-resolution methods in the field of face super-resolution and offers innovative applications in many areas such as biometric security, digital media, forensics and surveillance systems. The model produces more realistic and natural results by taking into account not only local details but also global structural integrity. The model is expected to make a significant contribution especially in areas such as security systems, enhancement of old photographs, clarification of poor quality security camera images and biometric verification. Various enhancements are possible to further improve the performance of the model. An adaptive adjustment of the learning rate could allow the model to perform optimally for different datasets. In addition, integrating the model into a hybrid structure with different deep learning techniques such as generative adversarial networks (GAN) and transformer-based models can further improve the super-resolution performance. Furthermore, tests on larger datasets can be conducted to increase the generalization capability of the model and make it more suitable for real-world applications. In conclusion, the proposed work proposes a model based on deep reinforcement learning and residual dense blocks within residuals (RRDB) to solve the face super-resolution problem. With a strategic part selection mechanism, the most critical regions of the face image are identified and the details are iteratively improved while maintaining the global structure integrity. Considering the wide range of applications and superior performance of the model, it can be said that it offers a significant advance in the field of facial super-resolution and image enhancement.
Author
Dr. Emre Altınkaya
Institution
Sakarya University
Elektrik Mühendisliği Bilim Dalı
How to Cite
Emre Altınkaya (Doctorate thesis). High resolution facial image acquisition with reinforcement learning, 2025, Sakarya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)