Determination of exposure using statistical data in machine learning: A case study in Sakarya
2024
0 views
0 downloads
Advisor: Prof. Dr. Naci Çağlar
Abstract (EN)
In the disaster management cycle, completing pre disaster emergency response and aid preparations and making financial investments are especially important to minimize human and monetary losses. For this reason, to effectively complete the preparation phase, which is the first phase of the disaster cycle, it is necessary to analyse losses in the pre-disaster period and take precautions accordingly. Studies conducted for loss analysis are based on three basic components: hazard, exposure, and vulnerability components. Among the loss components, the phase that requires the most time and workforce is the creation of the exposure model, which is defined as the characterization of the values that will be exposed to hazards and the determination of their spatial distribution. One of the biggest uncertainties in current exposure model is that the smallest resolution used for spatial distribution is at the provincial level due to lack of data. Another deficiency is the insufficient of building stock information in the data sets used. The first source of data used for exposure models to be created at the national level is building permit statistics collected by the country's central statistics office and additional information collected population censuses. However, the data may not contain sufficient quantity or detail information for every nation. For this reason, uncertainties occur in the resolutions and information level of the models and additional information is needed. Within the scope of this thesis, a generalized building exposure model for all of Türkiye was created using various data sets, statistical summaries, expert opinions, statistical distributions and machine learning algorithms. The developed building exposure model is based on three main data sets. The first data set contains information about the physical characteristics, purpose of use, and occupancy status of each building within the borders of Sakarya province for use in urban renewal studies conducted by Sakarya Metropolitan Municipality (SBB) and Disaster and Emergency management Presidency (AFAD). This data set also includes features that can be used as labels in machine learning, such as plan irregularity, soft floor irregularity and building order, which are key factors in the vulnerability of building to seismic hazards. Apart from this feature, there are also important features such as the number of floors independent sections, load-bearing system types, infill wall types for masonry buildings and license date. As the second main data, there are permit statistics of buildings constructed between 1992 and 2023 in the districts within the borders of Türkiye, collected by the Turkish Statistical Institute (TUIK). In this data set, the building classification recommended by European Union the create common standards was used and the features are presented collected in these classes. Similar to SBB, permit statistics include physical characteristics such as the intended use of buildings, structural classes, number of floors, and number of flats. Unlike SBB data, it contains more summary building physical characteristics throughout Türkiye. Finally, the third main data used is the CORINE 2018 artificial land cover, which is carried out with the European Environment Agency and determines the land covers of xxii the member countries, including Türkiye, using satellite images. This data set contains polygons of residential areas at the district level. Within the scope of the thesis, the first two data sets were used for building characterization and the third data set was used for the spatial distribution of buildings. Although the SBB data set has a detailed feature space, certain statistical operations and transformations were carried out before use, as there are empty cells in the data set, and it only contains stock information for a certain region. The TUIK data set on the other hand, lacks the detail level of SBB data, but compared to SBB data, it includes building features in all districts within the period it represents. The biggest deficiency in the TUIK data set is that it completely lacks information on occupancy rates and buildings built before 1992. Within the scope of the thesis, these two date sets were combined using a common building characterization taxonomy to compensate for each other's deficiencies, and the missing information was completed with statistical methods and distribution within this complete data. K-Nearest Neighbour algorithm (KNN) has been used for missing data and beta distributions have been used for year distributions. The data on the interactive earthquake map published by the Natural Disaster Insurance Institution (DASK) for the exact building numbers for the provinces were used. The missing building before 1992 in the building permit data were completed with building property distributions before 1992 in the SBB data and buildings with construction years of 1992-1999 in the building permit data. Then in order to find the number of people living in a building class based on TUIK data, the results of the Address Based Population Registration System (ADNKS) at the district level, also share in TUIK, were obtained. District populations were obtained according to the flat rates in the districts and by using the household distribution according to the year of construction in the summary statistics of the Building Quality and Housing Survey (BKNA) published by TUIK. Finally, the reconstruction costs for each building were added by using the approximate unit costs of the building to be used in the calculation of architectural and engineering service fees published in the Official Gazette by the Ministry of Environment, Urbanization and Climate Change, and the usage type and surface area of each building. In this full data set created for Türkiye, 11,409,354 buildings were calculated and a reconstruction cost of $ 2.5 Trillion was calculated for the reconstruction of these buildings. In order to obtain more precise building fragility in exposure models, it is necessary to use features that may affect fragility in existing data. Within the scope of the thesis, machine learning models were trained to predict three building irregularities that are found in SBB data but not in TÜİK data. Random Forest Classification was used as the algorithm and "Irregularity in Plan", "Soft Floor" and "Building Order" in the SBB data were used as prediction labels. The accuracy rates of the models verified within the SBB data were obtained as 80.36% for the Irregularity in Plan label, 85.73% for the Soft Label and 71.93% for the Building Order. Then, prediction results were obtained on all data with these models and added to the exposure model. Finally, for the spatial distribution of the created data set, the district locations where the building is located were obtained. Studies have proven that the spatial distribution of exposure models has an impact on the uncertainty of the model. For this reason, representing all the buildings in the districts on a single point will increase the uncertainty of the model, so the spatial distribution of the buildings must be distributed using appropriate methods. In this thesis, artificial land covers defined on the CORINE data set were used for the spatial distribution of buildings. The artificial cover closest to the district locations was calculated and the buildings were distributed xxiii homogeneously on this cover. Finally, all data was added to the map of Türkiye along with their distribution.
Author
Dr. Muhammed Ali Haşıloğlu
Institution
How to Cite
Muhammed Ali Haşıloğlu (Master Thesis). Determination of exposure using statistical data in machine learning: A case study in Sakarya, 2024, Sakarya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)
