Master'sOpen Access

Detecting abnormal drinking water consumptions and developing forecast models by machine learning methods

2023
0 views
0 downloads
Advisor: Doç. Dr. İhsan Hakan Selvi

Abstract (EN)

In this study, it is predicted that there may be a certain order in the consumption of an important need such as drinking water by the household, as well as irregular consumption depending on different factors. Increasing population, limited drinking water resources, developing infrastructure and technology have increased the demand for drinking and utility water. There is a search for alternative water sources to meet this demand, but with this study, it is foreseen that these demands can be met by not wasting existing water and using it more efficiently. By using machine learning (ML) methods, which is a sub-branch of artificial intelligence (AI), drinking water consumption data in the past periods were analyzed, and ordinary and unusual consumption behavior models were extracted. It is envisaged that by detecting abnormal consumptions that may occur in drinking water consumption and informing the subscribers about this issue, it will be ensured that the consumption in the household remains within the normal consumption range. Although the amount of data collected, recorded and processed in today's IT world has increased significantly, it is known that the exact analysis is difficult in terms of time and cost. In this study, subscriber, meter, consumption, bill and payment data of 8,224 residential subscribers, whose water meter index reading is more than 160 periods throughout the province of Kayseri, between 2006 and 2022 (first 6 months) were taken into account. The data are combined on a spatial subscriber basis and a dataset which have 41-features is obtained. Prepared data set; It has been transformed into a 24-featured dataset using data preprocessing such as data integration, data transformation, noise identification and cleaning, and completion or cleaning of detected missing values. The scores of each of the features affecting drinking water consumption were calculated using basic statistical values and feature selection methods. It has been determined whether the features affect drinking water consumption, and if so, how much. In feature selections, the determination of the number of features sufficient to meet the purpose was determined by taking the ratio of the variance explained criterion (Variance Explained Criteria) as 90%. Thus, threshold values were determined that minimum 6 and maximum 11 features would be sufficient and that this number of features would remain in each feature selection method. Information gain (IG), gain ratio (GR), symmetrical uncertainty coefficient (SU), pearson correlation coefficient (r), f-score and random forest (RF) feature selection methods were used for the selection of features affecting consumption. Thus, 6 different sub-datasets were obtained over the determined threshold values. In addition, the 7th sub-dataset at least 3 times selected features which consisting of by the feature selection method was obtained. The number of meters exchanged for each subscriber, the number of contracted users, the days of consumption measurement, consumption amounts, invoice amounts, payment status data are likely to be different, and the number of features is also considered to be different. Therefore, sub-datasets consisting of the selected features were recorded in the study specifically for each subscriber. With all datasets, Tukey outlier labeling (TOL), isolation forest (IF), z-score, copula-based outlier detection (COPOD), median absolute deviation (MAD), local outlier factor (LOF), and elliptic envelope (EE) ML anomaly analysis methods have been used. Abnormal and normal drinking water consumptions were determined by using these 7 different ML anomaly analysis methods. As a result of the anomaly analysis, abnormal consumptions detected by each method were scored, and the abnormality score of each anomaly consumption behavior was calculated with the sum of the scores of 7 different anomaly methods. Each observation point detected by each algorithm is weighted with 1 point. Since 7 algorithms are used, a maximum of 7 points will be weighted for each observation point. According to the outlier scores, all consumption values were labeled with 4 different consumption classes (Normal, Caution, Risky, Extreme), and the dataset was made supervised. With the obtained supervised data sets, consumption class estimation models were developed using decision trees (DT), gaussian naive bayes (NB), k-nearest neighbors (KNN), logistic regression (LJR), multilayer perceptron neural network (MLP-NN), RF and gradient augmentation (GB) methods. Drinking water consumption prediction models are compared with ACC, R2 performance metrics, and the best prediction model is selected for each drinking water subscriber. Since the consumption data consists of TS data, the first years are education data and the last years are control data, and they are used by dividing them by 90% and 10%. The error matrices of all anomaly detection models were calculated and evaluated as a basis for the performance measurement between the predictive values and the actual values of the models detecting abnormal drinking water consumption classes. In addition, the developed models were compared with Accuracy, Sensitivity, Sensitivity, Sensitivity, Specificity, MAE and MSE performance metrics. Among the abnormal consumption detection models developed for each subscriber, the model with the highest accuracy was determined as the best model. Among the abnormal consumption detection models, those with the same Accuracy performance were compared according to their Sensitivity and Precision ratios, and the one with higher performance was preferred. A total of 921,088 models were developed, of which 16,077 showed maximum accuracy performance as the best model. The Accuracy performance ratios of the models were determined as the lowest 43%, the highest 100% and the average 85%, and it was understood that the models obtained by the EE method were mostly the best models, and the models obtained with the Z Score and LOF methods provided maximum Accuracy at very low rates. In the consumption class prediction models obtained in the study, the accuracy performance metric of 90% and above was preferred as the best model. Thus, it was seen that the best consumption class prediction models were obtained from the sub-datasets (V1, V2, V3, V4, V5, V6, V7) of 7,150 (86.94%) of 8,224 subscribers taken as a sample. It was observed that the best consumption class prediction models were obtained from the basic dataset (V0) of the remaining 1,074 subscribers (13.06%). The maximum Accuracy performance model developed for each subscriber was obtained by using the V5 sub-dataset obtained by the F-Score feature selection method at the highest rate. It has been understood that successful results have been obtained with other sub-datasets at rates close to each other. As a result of the study, it has been proven that abnormal drinking water consumptions can be determined by ML methods and consumption classes can be estimated by ML methods. In addition, with the feedbacks aimed at influencing individual consumption behaviors, it has been shown that drinking water can be used more efficiently without wasting, and a different perspective has been presented to the investment planning and management approaches of water managers. Thus, a model has been developed that can contribute to overcoming very high demands, detecting and tracking subscribers with potential for loss and illegal use, and raising social awareness for water saving. The main purpose of the study is to sensitize the subscribers and to reduce the abnormal consumptions in their consumption behaviors, in other words, to reduce them to normal levels. In this case, it is predicted that 139,427 m3 of water will be used with 20% savings in the consumptions in the 'Extreme' class, 154,058 m3 with 15% savings in the consumptions in the 'Risky' class, and 169,402 m3 of water with 10% savings in the consumptions in the 'Caution' class. Thus, a more efficient use of 462,887 m3 of 2.58% water in the system will be achieved. The study helped to meet the rapidly increasing urban population and water needs. Contributed to the delivery of equitable water services in a world facing increasing water scarcity and environmental degradation. It has been shown what factors can affect drinking and domestic water consumption. It has been shown that there is a regular behavior pattern and seasonality in water consumption. It has been determined that there are periods that disrupt the consumption order. It has been revealed that abnormal consumptions that disrupt the consumption pattern may be caused by meter reading error, lost and illegal consumption, meter measurement error and wasteful consumption. It has been shown that drinking water consumption behavior can be modeled with ML algorithms, which is a sub-branch of AI, and can be disciplined with feedback. The subject of the study is related to the consumption of drinking and utility water in households and a model has been developed that can be used in all service areas where mass consumption such as electricity, natural gas, internet, GSM is required and where a distribution network, subscriber management system and infrastructure is needed. In addition, these systems include many requirements such as consumption measurements and follow-up, quality controls, fault detection, maintenance and repair follow-up, inspection, cost-reducing measures. From this point of view, it is foreseen that the study can contribute to these areas as well.

Author

Dr. İsmail Güney

How to Cite

İsmail Güney (Master Thesis). Detecting abnormal drinking water consumptions and developing forecast models by machine learning methods, 2023, Sakarya University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Sakarya University