Part control with machine learning from data on CMM device
2024
0 views
0 downloads
Advisor: Doç. Dr. Nuray Canikoğlu
Abstract (EN)
In today's highly competitive manufacturing industry, quality control has become a pivotal function that ensures product reliability, consistency, and customer satisfaction. As product designs become more complex and tolerances more stringent, traditional quality inspection processes often fall short in meeting modern expectations for speed, accuracy, and objectivity. Manual inspection, which depends heavily on the expertise and judgment of quality control personnel, can result in subjective decisions, inconsistencies across shifts, and time-consuming evaluations. These limitations underscore the growing need for intelligent, automated systems that can accurately and efficiently assess part quality. This thesis proposes a novel machine learning-based quality control system that uses measurement data from Coordinate Measuring Machines (CMMs) to autonomously classify parts as either "ACCEPTED" or "REJECTED," thereby eliminating the need for human intervention in the decision making process. CMMs are widely utilized in dimensional metrology due to their ability to perform precise three-dimensional measurements on complex geometries. By using tactile or optical probes, CMMs can gather comprehensive data on dimensional and geometrical tolerances such as flatness, roundness, cylindricity, angularity, parallelism, and true position. Typically, these measurements are compared against CAD models or technical drawings that define geometric dimensioning and tolerancing (GD&T) limits. Although highly accurate, interpreting CMM data requires significant expertise, and manual evaluation of these measurements often results in inconsistencies, especially when measured values are close to tolerance limits. In this study, supervised machine learning algorithms were applied to historical CMM datasets that were labeled by experienced quality control engineers. The data consisted of measurement features extracted from two different part types (Part 1 and Part 2) that vary in geometry, complexity, and tolerancing requirements. These features serve as inputs to classification algorithms which predict whether each measured part should be accepted or rejected. The ultimate goal is to create a fully automated decision support system that can replicate expert-level classification accuracy without human oversight. Four machine learning algorithms were selected for model training and evaluation: Random Forest, Support Vector Machine (SVM), Naive Bayes, and K-Nearest Neighbors (KNN). Each algorithm represents a distinct category in the machine learning spectrum. Random Forest is an ensemble method that combines multiple decision trees to improve prediction accuracy and robustness. It is particularly well suited for handling non-linear data and noise. SVM is a powerful classifier that identifies the hyperplane that best separates the data into two categories, often performing well in high-dimensional spaces. Naive Bayes is a probabilistic model based on Bayes' theorem and assumes feature independence; it is known for its xxiii simplicity, speed, and good performance on small to medium-sized datasets. KNN, in contrast, is a non-parametric algorithm that assigns class labels based on the majority class of the nearest neighbors in the feature space, offering an intuitive yet computationally demanding approach to classification. The dataset for each part was preprocessed using normalization techniques to ensure uniform scaling of features. The data was then split into training and testing sets to evaluate model generalizability. Performance metrics included accuracy and macro averaged F1-score. The accuracy metric quantifies the proportion of total correct predictions, while F1-score considers both precision and recall, making it particularly important when the class distribution is imbalanced, as in the case of quality control where the "REJECTED" class often has fewer samples. Results for Part 1 indicated that the Random Forest algorithm achieved the highest classification performance, reaching 97.5% accuracy and 0.971 F1-score. These results suggest that Random Forest is highly effective at capturing the complex interactions among measurement features that determine part quality. Naive Bayes followed closely, achieving 90% accuracy and 0.898 F1-score, making it a strong candidate for fast and lightweight quality assessment systems. SVM delivered 82.5% accuracy and 0.799 F1-score, performing reasonably well but showing some vulnerability in correctly identifying rejected parts. KNN also achieved 82.5% accuracy but had the lowest F1-score (0.664), reflecting its struggle with class imbalance and inconsistent decision boundaries in high-dimensional spaces. For Part 2, which featured a different set of geometric and dimensional characteristics, the performance rankings shifted slightly. Naive Bayes emerged as the most effective classifier with 85% accuracy and 0.840 F1-score. Random Forest maintained strong performance but showed slightly lower F1-score (0.785), suggesting less balanced classification across the two output classes. SVM again showed moderate performance with 77.5% accuracy and 0.734 F1-score, while KNN continued to underperform with 80% accuracy and 0.724 F1-score. These results emphasize the necessity of selecting algorithms based on part-specific characteristics, data distributions, and class imbalances. Beyond performance metrics, the practical benefits of the proposed system are significant. The automation of quality decision-making reduces dependency on highly skilled operators, minimizes variability across shifts, and enables real-time feedback for process control. Moreover, the early identification of defective parts prevents downstream assembly issues, reduces scrap and rework costs, and enhances overall production efficiency. The system's adaptability also allows it to be retrained on new parts with minimal reconfiguration, making it highly scalable for diverse manufacturing environments. The developed system fits well into the framework of Industry 4.0 and smart manufacturing. By integrating machine learning into quality inspection, it contributes to the development of cyber-physical systems where physical measurement devices communicate with digital models and algorithms to support autonomous decision making. In future applications, this system can be enhanced by incorporating data from IoT-enabled sensors that capture real-time production parameters such as tool wear, cutting temperature, or vibration. Predictive quality control becomes feasible when this rich data is fed into time-aware models such as Recurrent Neural Networks (RNN) or Long Short-Term Memory networks (LSTM). Additionally, integration with cloud xxiv platforms can facilitate centralized monitoring, model updates, and real-time alerts across multiple production lines or factories. Comparative analysis with existing literature reveals that while machine vision systems and deep learning-based defect detection are common in quality control research, the use of metrology-grade CMM data for machine learning classification remains relatively underexplored. Most research focuses on image classification or simple sensor readings rather than high-precision dimensional data. This thesis fills that gap by proposing a comprehensive framework that connects physical measurements to intelligent classification logic. By doing so, it enables a new class of AI-driven quality inspection tools that can be deployed in precision manufacturing contexts such as aerospace, automotive, and medical device production. Furthermore, the system developed in this study offers a foundation for numerous future enhancements. Feature selection techniques such as Principal Component Analysis (PCA) or Recursive Feature Elimination (RFE) may be used to reduce dimensionality and further improve model efficiency. Imbalanced data handling methods such as SMOTE (Synthetic Minority Oversampling Technique) or cost sensitive learning can be integrated to increase robustness, particularly for rare rejection cases. Advanced evaluation metrics, including ROC curves, AUC scores, and Matthews Correlation Coefficient (MCC), can be added to deepen performance analysis. Additionally, explainable AI (XAI) techniques such as SHAP or LIME could be employed to increase trust and interpretability of model decisions for process engineers. In conclusion, this thesis demonstrates the viability and effectiveness of using machine learning algorithms on CMM measurement data to develop an automated quality control system. The system not only achieves high levels of accuracy and class balance but also provides a scalable, adaptable, and operator-independent solution for real-time quality assurance. As digital manufacturing technologies continue to evolve, such intelligent systems are expected to become an integral part of quality management strategies, offering enhanced precision, reliability, and productivity.
Author
Dr. Zeynep Birgin
Institution
How to Cite
Zeynep Birgin (Master Thesis). Part control with machine learning from data on CMM device, 2024, Sakarya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)
