Learning object-centric representations based on slots in real world scenarios
2025
0 views
0 downloads
Advisor: Prof. Dr. Yücel Yemez
Abstract (EN)
A central goal in artificial intelligence is to enable machines to perceive the visual world as a composition of distinct objects. This ability for object-centric understanding is essential for generative models that support fine-grained, controllable content creation and editing. However, state-of-the-art diffusion models process images holistically and are conditioned on text, creating a semantic misalignment when tasked with object-level manipulation. As a result, researchers face a fundamental challenge: either adapt powerful but text-biased models or build specialized models from scratch, often with reduced capacity. This dissertation addresses this problem by introducing a framework that adapts pretrained generative models for object-centric image and video synthesis. Our analysis highlights a core challenge in current approaches: achieving high-quality generation requires balancing global scene coherence with disentangled, object-level control. To address this, we propose an adaptation strategy that integrates object-specific conditioning into pretrained models while preserving their valuable priors. Extending this framework to video further amplifies the difficulty, as maintaining temporal coherence and consistent object identity across frames is critical. For static images, we introduce SlotAdapt, a method that augments diffusion models with lightweight slot-based modules. A register token captures background and style, while slot-conditioned components encode object-specific information. This dual-pathway design mitigates text-conditioning bias and provides precise, object-centric control, leading to state-of-the-art results in object discovery, segmentation, compositional editing, and controllable image generation. We then extend the framework to video. Using Invariant Slot Attention (ISA) to disentangle object identity from pose, combined with a Transformer-based temporal aggregator, our approach ensures consistent object representation and dynamics across time. This framework sets new benchmarks in unsupervised video object segmentation and reconstruction, while enabling advanced video editing capabilities, including object removal, replacement, and insertion, all without explicit supervision. Overall, this work establishes a general and scalable approach to object-centric generative modeling for both images and videos. Beyond setting new technical baselines, it expands the design space for interactive and controllable generative tools, bridging the gap between human object-based perception and machine learning models. These contributions open new directions for structured, intuitive, and user-driven AI applications in creative, scientific, and practical domains.
Author
Dr. Adil Kaan Akan
Institution

Koç University
Bilgisayar Bilimi ve Mühendisliği Bilim Dalı
How to Cite
Adil Kaan Akan (Doctorate thesis). Learning object-centric representations based on slots in real world scenarios, 2025, Koç University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Koç University
- Obje tabanlı akıl danışma-tavsiye iletişimi tasarımına ilham kaynağı olarak Türk kahve falı(2017)
- Ekom-Eczacıbaşı'nın Rusya piyasasındaki pazarlama stratejileri(1995)
- Barok döneminde Balkanlar Osmanlı Avrupası'nda mimaride, dekorasyonda, himaye ve kültürel üretim modellerinde dönüşüm, 1718-1856(2006)
- De Rham-Witt kompleks(2011)
- Erteleme kısıtlı tek makine çizelgeleme(2014)
- Sarayda Osmanlı tütsüleme gelenekleri: Topkapı Sarayı buhurdanları(2015)