Modelling driver-related traffic accidents through data mining approach: The case of city of Sakarya
2024
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Hakan Aslan ; Prof. Dr. Nilüfer Yurtay
Abstract (EN)
Transportation is the purposed movement of people, goods, and animals from one place to another. Transportation types are classified as road, rail, sea, and air. In our country, the most widely used type of them has been road transportation for years. The heavy traffic on the roads, coupled with the increase in private vehicle usage, makes traffic safety as a challenging issue due to the increases in the risk of collisions. Traffic accidents are a major issue affecting human life directly and indirectly, both globally and in Türkiye. The United Nations aims to reduce traffic accident-related injuries and deaths by 50% by 2030 and to eliminate them entirely by 2050. Responsible institutions for transportation in Türkiye (General Directorate of Highways, General Directorate of Security, The Gendarmerie General Command, Municipalities, Ministry of National Education) have various practices related to transportation planning, traffic management, and traffic safety. However, achieving the "vision of zero fatalities" still requires numerous measures as evident from recent accident statistics. Traffic accidents, generally resulting from the combination of two or more factors, cause multidimensional damages. According to 2022 data in Türkiye, traffic accidents lead to a daily average of 14 fatalities, 790 injuries, and 2837 property-damage-only incidents, posing a significant public health problem. Loss of life, post-accident physical disabilities, and property losses significantly impact both the country's economy and the social life of the people leading to safety concerns on the roads and psychological distress for those involved in accidents and their families. Despite initiatives in recent years such as the design of "forgiving roads" to tolerate user errors, there are still many aspects that need improvement as far as road safety is concerned, as indicated by recent accident statistics. Errors leading to traffic accidents are categorized into three main related groups: human (driver, pedestrian, passenger), vehicle, and road. The most crucial of these errors is the human related factors, ranging between 97-99% from 2011 to 2022, with the driver factor accounting for 87-90%, as stated in a report prepared by the KGM in 2022. Türkiye with the population of 86,907,000.00 has 259,072 km-long highways. The number of vehicles registered in traffic is 28,183,745.00 among them 14,967,044.00 being automobiles according to the 2023 records. The average numbers of cars per 1000 people for Türkiye and EU countries are 167 and 560, respectively. While the number of deaths per 1 million cars in Türkiye is 366 people, the related figure for EU countries is 76 reflecting the need to make significant progress in road safety engineering and applications. This study aims to examine fatal and injury accidents in the city of Sakarya using data mining analysis methods. Two different analyses were conducted in the study. The first involved exploring possible relationships among traffic-related risk factors using association rules. The second used decision trees to predict the severity of accidents involving only cars. The study identified which conditions co-occur most frequently in the traffic accident dataset and which elements (classes within attributes) were more prevalent overall. Decision trees are widely used for classification in situations like traffic accidents. In decision trees, while each branch departing from a node answers a question in that area, each leaf represents a decision outcome. The selection of the root node, branching and pruning decisions, and the subsequent classification depend on the algorithm used when building the decision tree. In this study, the traffic accident dataset was examined, and it was concluded that accident severity could be classified. Among the four attributes related to the accident outcome in the dataset, one is the outcome attribute containing two classes: fatal and non-fatal injuries. Being only 0.9% of the accidents fatal makes predictions for the remaining injuries is challenging. The other attributes related to the accident outcome (treated as separate attributes) are associated with drivers, passengers, and pedestrians. Since numerous accident data sets do not involve information regarding passengers and pedestrians, drivers are regarded as the most affected and statistically contributing the most to the accidents. Therefore, in this study, the driver's accident outcomes were predicted using decision trees, a classification method. Some of the decision tree algorithms are CHAID (Chi Square Automatic Interaction Detection), CART (Classification and Regression Trees), C 4.5, ID3 and Random Tree. Since the dependent and independent variables are categorical (nominal), the CHAID algorithm was deemed suitable within the scope of this thesis. As the purpose of the second analysis in this thesis is to classify the characteristics of the faulty drivers involved in the accidents, and since the vehicle-size attribute makes it difficult to fulfill this purpose, the vehicle with the largest weight in the dataset (55%) was selected as automobiles. The reason why this objective is challenging is that there are studies suggesting that vehicle size affects injury severity. It has been reported in the literature that the most closely related vehicle characteristic to injury severity is the size of the vehicle. Although there is a high correlation between weight and length, weight is reported to be more common in studies. As far as the analysis is concerned, chi-square values were, first, calculated for all attributes, and then, due to the dominant effect of some attributes obscuring driver related parameters, it was decided to exclude these attributes from the dataset. With this regard, chi-square values between the remaining and the output attributes (driver's accident outcome) were calculated and determined. The attribute with the greatest relationship magnitude at the root node of the tree, indicating the largest relationship with the dependent attribute, was obtained as driver error. Data set is divided into two parts to create the decision trees of the model; 70% of which as training and 30% as test data. This distinction was made linearly and as the accidents in the data set were recorded and presented according to the date/year on/in which they occurred, the first 5 years of the available data have been taken as education, and the last 2 years as testing data. To create a CHAID decision tree and classify the driver related severity of the accidents, chi-square values were first calculated with the dataset containing all attributes. Subsequently, due to the dominant influence of certain attributes leading to the concealment of driver-related situations, it was decided to remove these dominant attributes from the dataset. Remaining attributes were then used to calculate chi-square xxix values between the output attribute (dependent variable), which is the driver's posed severity level of the accident, and the remaining attributes. According to the magnitude of the relationship between the driver's fault attribute and the dependent attribute, the following attributes were ranked: driver under the influence of alcohol, road section, time zone, driver's age, presence of driver's license, age of the automobile, driver's education level, and daylight. According to the decision tree created, drivers using alcohol were all injured, and drivers closely following the preceding vehicle had their accident outcomes classified according to the driver's age group. The outcomes of drivers who did not adhere to turning rules varied according to their genders, and the accident outcomes of drivers who did not adjust their speed to the road and weather conditions primarily depended on the road section. In the next stage, the casualties were observed to change according to the gender for 3 and 4-way intersections. Among male drivers, they varied according to age group. In roundabouts, on the other hand, while the driver's accident outcome changed according to the time zone, at intersections and crossings, the outcomes of drivers not yielding through the right-of way rule changed according to the driver's education level. Among high school graduates, the outcome also varied according to the age of the vehicle they used. In accidents where the driver's fault was improper lane changing, the educational level for accidents involving elementary school graduates played an important role, and within this group, the results varied according to the daylight conditions and the age of the vehicle used.
Author
Dr. Zeliha Çağla Kuyumcu
Institution
How to Cite
Zeliha Çağla Kuyumcu (Doctorate thesis). Modelling driver-related traffic accidents through data mining approach: The case of city of Sakarya, 2024, Sakarya University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Turkey according to the records of the House of Commons (1918-1922)(2011)
- The effect of digital accounting applications on preventing accounting errors and frauds: A research on professional members(2025)
- Mawlana Yaqub-i Charkhi And His tafsir(2024)
- Investigation of friction and wear behaviors of Al/AlB2 composites materials(2012)
- Extraordinary events in the commentary named Rûhu?l-Beyân of İsmail Hakkı Bursevî(2008)
- Karl R. Popper`s democracy approach(1998)
