Random Forest-Based Classification for Early Detection of Student Dropout
Keywords:
Student Dropout, Random Forest, Classification, Machine Learning, Educational Data MiningAbstract
Student dropout is one of the major challenges in higher education because it affects academic performance, graduation rates, and institutional reputation. Early identification of students who are at risk of dropping out is important so that universities can provide appropriate intervention strategies and academic support systems. Along with the growth of educational data, machine learning techniques have become increasingly useful for analyzing student behavior patterns and predicting academic outcomes more effectively. This study aims to classify student dropout using the Random Forest algorithm on the Predict Students' Dropout and Academic Success dataset obtained from the UCI Machine Learning Repository. The research process includes data preprocessing, feature selection, model training, and performance evaluation. Random Forest is used to classify students into dropout, enrolled, or graduate categories based on academic, demographic, and socioeconomic attributes. Model performance is evaluated using accuracy, precision, recall, F1-score, and confusion matrix analysis. The results of this study are expected to support universities in identifying students at risk of dropout and improving student retention strategies through data-driven decision-making. In addition, this research contributes to the application of machine learning techniques in educational data mining for academic prediction systems.
