DoctorateOpen Access

Text analysis of standard construction contract documents by the application of text mining and machine learning techniques

2025
0 views
0 downloads
Advisor: Doç. Dr. Latif Onur Uğur

Abstract (EN)

In this dissertation, a holistic classification framework based on natural language processing (NLP), text mining, and machine learning (ML) techniques has been developed to make sense of the multidimensional content structure of standard Design-Build contracts. Focusing on the Design-Build project delivery method, the research was conducted on the standard English contract texts of the 'FIDIC Conditions of Contract for Design, Build and Operate' and the 'JCT Design and Build Contract'. The contract provisions were classified along both thematic (obligations, tasks, discretionary expressions) and functional (cost, time, quality) dimensions. The study is built upon a three-stage experimental design. In the first stage, contract provisions were categorized into thematic classes such as "obligations," "to do," "optionals," and "other general statements". Classical text representation methods (TF-IDF, BoW, Word2Vec: CBoW & Skip-Gram) were employed in combination with fundamental Machine Learning (ML) algorithms such as "Support Vector Machines (SVM), Logistic Regression (LR), Naive Bayes (NB), K-Nearest Neighbors (KNN), and Decision Trees (DT)", as well as Ensemble Learning (EL) methods including "XGBoost, AdaBoost, Hard Voting, Soft Voting, and Stacking". The combination of CBoW representation and XGBoost yielded the highest performance with 96.6% accuracy (Acc) and 96.2% F1 score, demonstrating results competitive with the literature. In the second stage, the texts were labeled along the dimensions of "cost," "time," and "quality," and the classification of these managerial themes was targeted using the same technical framework. Although the Skip-Gram + XGBoost combination achieved an F1 score of 88.6%, the intertwinement of cost and time expressions introduced partial uncertainties in model performance. This highlighted the need for more robust contextual analysis tools. Building on this, in the third stage, four different transformer-based language models (BERT, alBERT, RoBERTa, DistilBERT) were integrated with sequential Deep Learning (DL) architectures such as "GRU, LSTM, and RNN". In two separate multi-class classification tasks, the RoBERTa + LSTM combination achieved superior performance with 98.06% Acc, clearly demonstrating the success of DL models in tasks requiring contextual semantic distinction. The findings indicate that artificial intelligence-based approaches have a high potential for the automatic, contextual, and reliable analysis of standard construction contracts. The model developed within the scope of this dissertation is capable of producing meaningful classifications across multi-document structures; in this respect, it holds applicability for decision-support processes such as contract management, bid analysis, and risk assessment. Comprehensive comparisons of contextual models sparsely addressed in the literature have been conducted, thereby contributing to the digital transformation goals specific to the construction industry. This allows project managers and legal professionals to interact with contract texts, receive timely feedback, and enhance the efficiency of their decision-making processes.

Author

Anıl Demircan

How to Cite

Anıl Demircan (Doctorate thesis). Text analysis of standard construction contract documents by the application of text mining and machine learning techniques, 2025, Düzce University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Düzce University