Master'sOpen Access

Efficient autonomous driving with foundation models

2025
0 views
0 downloads
Advisor: Assist. Prof. Dr. Fatma Güney

Abstract (EN)

Driving safely in urban environments is a task that requires a deep understanding of complex, dynamic environments and the ability to react swiftly and effectively to unexpected events. Recent advances in large foundation models have shown impressive capabilities to reason and generalize. In this thesis, we explore the potential and feasibility of foundation models for driving. First, we start by building on top of previous work showing promising performance in robotics tasks by formulating reinforcement learning as a language modeling problem using auto-regressive transformers. These approaches, however, assume the state is relatively simple, often representable as a single vector. In driving, the state is more complex, involving multiple agents that interact with each other, raising the need to find a compact, effective state representation. In Carformer, we propose a language model using learned, self-supervised, object-centric representation based on slot attention. We analyze different possible state representations in a privileged setting and find that our proposed representation outperforms both scene-level and hand-crafted object-centric representations, achieving the state-of-the-art on the Longest6 benchmark. After showing strong results with a language modeling approach in a privileged setting, we then focus on a more realistic, end-to-end setting. While foundation models further improve with scale, this results in additional inference latency. Latency is one of the main hurdles facing the deployment of large foundation models for real-time applications like driving. To address this, we propose a dual approach pipeline, ETA, that utilizes a large foundation model in tandem with a smaller, lightweight model. To minimize latency, we shift the computation of the large foundation model to an earlier time-step following it with a forecasting module to adapt the features to the present. Unlike previous work on dual approaches, we batch the large model inference to enable the decisions to benefit from the large model features at every time-step. Outperforming all previous approaches at a fraction of the latency, we highlight the feasibility of large foundation models in self-driving.

Author

Dr. Shadı Sameh Mohammed Hamdan

How to Cite

Shadı Sameh Mohammed Hamdan (Master Thesis). Efficient autonomous driving with foundation models, 2025, Koç University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Koç University