Detection of code similarities using abstract syntax trees and deep learning methods
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
2023
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Ali Nizam
Abstract (EN)
The aim of this thesis work is to design a system that creates a dataset indicating similarities in the source code based on the C# programming language, detects these similarities, and identifies points that need improvement. In the software world, understanding the structure of programs and detecting code similarities is of vital importance for a range of significant applications. These applications include code clone detection, software fraud monitoring, and software quality control. In this context, Abstract Syntax Tree (AST) based code similarity analysis offers a series of effective solutions. AST represents the syntactic structure of a program in a hierarchical way. ASTs can be used to determine the deeper semantic similarities of the code, as they reveal the structural features and details of the code. Therefore, AST-based analyses are powerful and flexible in determining code similarities. This study examines the basic principles, techniques, and applications of AST-based code similarity analysis. It also covers details on how ASTs are formed and used. In the work carried out, unlike other studies, a model detecting the similarities of sub-breakdowns of the code, not just based on classes or methods, has been put forward. Although it has been developed with various methods and tools, it was observed that there was no embedded library (embedding vocabulary) containing only ASTs, and therefore a vocabulary containing only relevant ASTs was created. The effects on code similarity are also examined by using triplet loss deep learning network methods, which are mostly applied on the similarities of images. For this purpose, a different approach has been introduced to detect code similarity using the triplet loss deep learning network by creating relevant data sets based on the ASTs of similar code and dissimilar code blocks. With a separate application, method, and block (if, for, while) based ASTs were created based on the codes of some sorting algorithms (Quick Sort, Bubble Sort, etc.). The line ranges of the relevant code blocks and the similarities of each AST line with other AST lines were extracted and a dataset was created. When the code similarities were measured with the cosine similarity method, based on the highest of these code similarities, an accuracy rate of 61.4% was achieved. The developed model was also compared with the tools used today for code similarity. The high accuracy rate achieved by using triplet loss has been evaluated as an important development in the practical use of the proposed technique.
Author
Necmettin Elmascı
Institution
How to Cite
Necmettin Elmascı (Master Thesis). Detection of code similarities using abstract syntax trees and deep learning methods, 2023, Fatih Sultan Mehmet Foundation University .
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fatih Sultan Mehmet Foundation University
- Examination of the relationship between antisocial behavior psychopatic tendencies and childhood mental trauma experiences in the adult male offenders convicted of physical and sexual violence(2025)
- The role of positive and negative emotion severity, anger symptoms, and depression level in the relationship between chronic pain and childhood traumas(2025)
- ربقيق قسم البالغة من سلطوط "شرح مفتاح العلوـ"ِِّّٓفِّ اخلوارزمي حلساـ الدِّين ادلؤذَّ(من أولو إ ُف آخر اعتبارات ادلسند إليو)(2025)
- Examining the effectiveness of "Emotion Socialization Training Program" on parents and preschool teachers of 48-72 month-old children(2025)
- Examination of the relation between styles of adult attachments and early maladaptive schemas(2025)
- القضايا الفقهية المعاصرة المتعلقة بالأحوال الشخصية في لصومال(2025)