DoctorateOpen Access

Learning object-centric representations based on slots in real world scenarios

2025
0 views
0 downloads
Advisor: Prof. Dr. Yücel Yemez

Abstract (EN)

A central goal in artificial intelligence is to enable machines to perceive the visual world as a composition of distinct objects. This ability for object-centric understanding is essential for generative models that support fine-grained, controllable content creation and editing. However, state-of-the-art diffusion models process images holistically and are conditioned on text, creating a semantic misalignment when tasked with object-level manipulation. As a result, researchers face a fundamental challenge: either adapt powerful but text-biased models or build specialized models from scratch, often with reduced capacity. This dissertation addresses this problem by introducing a framework that adapts pretrained generative models for object-centric image and video synthesis. Our analysis highlights a core challenge in current approaches: achieving high-quality generation requires balancing global scene coherence with disentangled, object-level control. To address this, we propose an adaptation strategy that integrates object-specific conditioning into pretrained models while preserving their valuable priors. Extending this framework to video further amplifies the difficulty, as maintaining temporal coherence and consistent object identity across frames is critical. For static images, we introduce SlotAdapt, a method that augments diffusion models with lightweight slot-based modules. A register token captures background and style, while slot-conditioned components encode object-specific information. This dual-pathway design mitigates text-conditioning bias and provides precise, object-centric control, leading to state-of-the-art results in object discovery, segmentation, compositional editing, and controllable image generation. We then extend the framework to video. Using Invariant Slot Attention (ISA) to disentangle object identity from pose, combined with a Transformer-based temporal aggregator, our approach ensures consistent object representation and dynamics across time. This framework sets new benchmarks in unsupervised video object segmentation and reconstruction, while enabling advanced video editing capabilities, including object removal, replacement, and insertion, all without explicit supervision. Overall, this work establishes a general and scalable approach to object-centric generative modeling for both images and videos. Beyond setting new technical baselines, it expands the design space for interactive and controllable generative tools, bridging the gap between human object-based perception and machine learning models. These contributions open new directions for structured, intuitive, and user-driven AI applications in creative, scientific, and practical domains.

Author

Dr. Adil Kaan Akan

How to Cite

Adil Kaan Akan (Doctorate thesis). Learning object-centric representations based on slots in real world scenarios, 2025, Koç University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University