Optimizing learned image compression models for complexity and rate-distortion-perception performance
2023
1 views
0 downloads
Advisor: Prof. Dr. Ahmet Murat Tekalp
Abstract (EN)
Lately, the rate-distortion performance of learned image compression models has surpassed that of traditional codecs by the virtue of recent advancements in learned entropy and context models. However, state-of-the-art learned models currently exhibit higher complexity and slower processing times compared to conventional image codecs. Furthermore, optimization of models for just rate-distortion performance as currently done does not result in the best perceptual image quality. This thesis addresses these issues. One of the contributions of this thesis is to explore the impact of the activation function on the performance of image compression, considering both objective and subjective evaluation criteria, as well as runtime efficiency. The widely used generalized divisive normalization (GDN) activation function is one of the reasons for its high complexity. Our findings reveal that the latent variables generated by hard shrinkage activation align more closely with a Laplacian distribution. Our method achieves comparable rate-distortion results, along with superior visual performance, at reduced computational complexity. The second contribution of this thesis lies in the exploration of practical approaches to the optimization of rate-distortion-perception (RDP) performance. To date, the use of mean squared error (MSE) in rate-distortion optimization (RDO) has remained the standard practice in the field of image and video compression. This has been beneficial for gauging codec performance by offering a quantitative measurement of results through peak-signal-to-noise ratio (PSNR). However, it's broadly accepted that PSNR does not accurately reflect the perceptual quality of images, making RDO unsuitable for codec optimization in terms of perceptual quality. Recently, the notion of RDP has been formally defined by Blau and Michaeli [1]. Yet, there's still a lack of practical methodology for setting the RDP function at a desired level in a feasible way. We propose a practical method to enable perception-distortion analysis by keeping the rate constant. This approach allows for a principled perceptual evaluation of the codec at predetermined bitrates. Additionally, we present a method for compressing a set of images at a desired RDP point by converting the problem to an integer linear programming model. Our experimental results provide essential insights into the practical analysis of RDP in learned image compression.
Author
Dr. Ogün Kırmemiş
Institution

Koç University
Elektrik Elektronik Mühendisliği Bilim Dalı
How to Cite
Ogün Kırmemiş (Doctorate thesis). Optimizing learned image compression models for complexity and rate-distortion-perception performance, 2023, Koç University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- De Rham-Witt kompleks(2011)
- Erteleme kısıtlı tek makine çizelgeleme(2014)
- Sarayda Osmanlı tütsüleme gelenekleri: Topkapı Sarayı buhurdanları(2015)