DIGITAL LIBRARY
A COMPARISON OF CLASSICAL, ENSEMBLE, AND DEEP LEARNING MODELS TO CLASSIFY STUDENT DROPOUT IN ENGINEERING PROGRAMS
Universidade Federal do Ceará (BRAZIL)
About this paper:
Appears in: EDULEARN26 Proceedings
Publication year: 2026
Article: 2599
ISBN: 978-84-09-88444-5
ISSN: 2340-1117
doi: 10.21125/edulearn.2026.2599
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
High dropout rates in Brazilian public universities represent a significant waste of resources and talent. To address this challenge, this study applies Educational Data Mining techniques to classify student dropout in Engineering programs at the Federal University of Ceará. We utilized a dataset comprising academic records from 2015 to 2022, a timeframe strategically selected to encompass the maximum allowable period for a student to complete their degree. Our methodology employs Mutual Information for feature selection to capture non-linear dependencies within sparse academic records. Moving beyond standard single-split evaluations, we utilized Stratified K-Fold Cross-Validation and Grid Search to prevent overfitting and ensure hyperparameter tuning across four distinct architectures: Support Vector Machines (SVM), Multilayer Perceptron (MLP), Random Forest, and LightGBM, along with a statistical validity of the comparative analysis using 5x2cv paired t-test. LightGBM emerged as the superior model, achieving a validation accuracy of 93% and 86% recall regarding the graduated class, with minimal variance. Crucially, hypothesis testing revealed that tree-based ensemble models demonstrated clear statistical superiority over both classical and deep learning approaches. Furthermore, the tests showed a significant difference favoring classical methods over the MLP, demonstrating that computationally heavy neural networks provide no classification advantage over well-tuned classical baseline algorithms for this specific tabular dataset. Ultimately, this framework offers a tool for academic coordinators to identify at-risk students and efficiently allocate pedagogical interventions.
Keywords:
Educational Data Mining, Dropout Classification, Ensemble Models.