DoctorateOpen Access

Perceptual video restoration: Task-specific versus pre-trained large generative models

2025
0 views
0 downloads
Advisor: Prof. Dr. Ahmet Murat Tekalp

Abstract (EN)

Video restoration problems are ill-posed inverse problems with a very large null space. Perceptual video restoration refers to choosing a feasible solution that is perceptually pleasing and temporally coherent. This thesis compares two paradigms for perceptual video restoration: supervised task-specific generative models and pre-trained large generative models applied in a zero-shot manner with some inference-time adaptation across diverse video restoration tasks. In the first part, we present a novel task-specific Video Super-Resolution (VSR) model trained under a perception-aware framework, which trades fidelity for spatio-temporal perceptual quality. Our architecture employs two discriminators: one that focuses on the realism of spatial texture, and the other on the perceptual naturalness of motion. This design is supported by a hypothesis called perceptual straightening of natural videos, which states that natural image sequences that appear curved in the intensity domain tend to follow straighter trajectories in the perceptual domain of human visual system. Extensive experiments demonstrate that this model achieves superior spatio-temporal perceptual quality compared to baseline approaches. In the second part, we explore a zero-shot, training-free paradigm to adapt large pre-trained image-based latent diffusion models (LDMs), for general-purpose video restoration tasks including video super-resolution, inpainting, deblurring, motion blur removal, and a combination of aforementioned tasks. While these models are rather flexible and offer quite good spatial generative capabilities, being inherently image-based, suffer from temporal inconsistencies when applied frame by frame, and mechanisms for enforcing temporal coherence remain relatively underexplored in this zero-shot setting. To address this, we introduce three contributions: (1) a cross-frame attention approach to foster feature propagation without any motion estimation; (2) perceptual straightness regularization to guide sampling in finding perceptually smoother temporal paths, and, (3) a multi-path ensemble sampling scheme that improves overall quality—including temporal consistency—during the diffusion denoising process. Quantitative and qualitative results reveal a trade-off between task-specific models and general zero-shot large-scale models: task-specific models achieve the best fidelity and perceptual scores when test degradations match training, but their performance drops under degradation kernel or dataset shift, where the zero-shot model remains competitive or superior and provides a flexible, training-free option—albeit with higher runtime. This thesis presents a detailed exploring of both paradigms, with a focus on the different approaches employed to address the challenge of temporal consistency—leveraging explicit training-based mechanisms in task-specific models and introducing inference-time temporal consistency techniques within the zero-shot diffusion framework.

Author

Dr. Nasrın Rahımı

How to Cite

Nasrın Rahımı (Doctorate thesis). Perceptual video restoration: Task-specific versus pre-trained large generative models, 2025, Koç University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University