Master'sOpen Access

Sentetik katılımcılar: büyük dil modellerinin insan anket katılımcılarına alternatif olarak değerlendirilmesi

2025
0 views
0 downloads
Advisor: Doç. Dr. Enis Kayış

Abstract (EN)

This thesis evaluates Large Language Models (LLMs) as "synthetic participants" in opinion surveys, addressing rising costs and biases in traditional data collection. Using 10,000 profiles from Wave 7 of the World Values Survey, ten LLM architectures were prompted to respond to seven trust-related questions. Performance was assessed via exact-match accuracy, mean absolute error (MAE), distributional divergence, and feature importance alignment, comparing LLM outputs to actual human participants and supervised machine learning models. Findings show that LLMs achieve a predictive accuracy that outperforms chance but lags behind specifically-trained supervised models. While errors were directionally accurate, divergence tests revealed a consistent "majority class fixation," where LLMs overrepresented dominant views and suppressed minority perspectives. Furthermore, while both machine learning models trained on human responses and those trained on LLM outputs rely on macro-institutional indicators, the first group is further shaped by cultural identity (ethnicity, language), whereas the second group relies on socio-economic status (income, social class). Consequently, LLMs function as "global average smoothers"—capturing broad trends but failing to replicate the heterogeneity required for high-fidelity social research. While promising for exploratory studies, LLMs remain an inappropriate substitute for human data in contexts requiring representational accuracy for minority or cross-cultural viewpoints.

Author

Dr. Alp Uçar

How to Cite

Alp Uçar (Master Thesis). Sentetik katılımcılar: büyük dil modellerinin insan anket katılımcılarına alternatif olarak değerlendirilmesi, 2025, Özyegin University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Özyegin University