TOWARD SYSTEMATIC GUIDELINES FOR THE DEVELOPMENT OF STUDENT DROPOUT PREDICTION MODELS
1 University of Cruz Alta (BRAZIL)
2 University of São Paulo (BRAZIL)
3 Pontifícia Universidade Católica do Paraná (BRAZIL)
4 Universidade Regional do Noroeste do Estado do Rio Grande do Sul (BRAZIL)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Student dropout in higher education is one of the most critical phenomena in the academic trajectory, with negative impacts on both students and educational institutions. The use of educational data mining techniques has proven to be a promising approach to mitigating this problem, enabling the construction of predictive models that identify students at risk of dropping out. However, the literature indicates that the construction of these models frequently lacks a structured approach. The selection of dropout indicators tends to be based on researcher or educator intuition rather than solid theoretical frameworks, which can result in inaccurate or incomplete diagnoses of the phenomenon's causes.
This work addresses this gap by proposing a set of systematic guidelines to steer the development of dropout prediction models, structured into three main phases: planning, execution, and documentation. The planning phase focuses on a clear definition of the problem and the well-founded selection of variables. The execution phase encompasses data processing and the application of machine learning algorithms. Finally, the documentation phase aims to ensure the reproducibility and transparency of the generated models. The objective of the proposed guidelines is to mitigate the fragmented and intuitive nature of building these models. It seeks to transform the creation process into a systematic, methodological activity, ensuring that the choice of indicators and algorithms is not arbitrary or limited solely to immediate data availability, but is guided by predefined criteria.
To evaluate the feasibility and quality of the guidelines, a preliminary validation was conducted through a content inspection. This stage was conducted by experts with solid experience in predictive models and the educational domain, who systematically analyzed the proposed guidelines. During the inspection process, the experts evaluated the completeness and correctness of each suggested guideline. The results of this inspection identified 15 problems in the proposed guidelines: 9 related to completeness (e.g., omissions of important actions) and 6 related to correctness (e.g., conceptual or procedural errors). It was noted that most of these problems occurred during execution, suggesting that the technical details of data modeling require further refinement.
Although some issues were identified, the inspection process validated the set of guidelines. Expert feedback supported targeted adjustments to the clarity and completeness of the guidelines, resulting in a consolidated version that systematizes field knowledge and is ready to be adopted as a support resource for predictive model development.
The main contributions of this work lie in the systematization of knowledge for dropout prediction, filling a methodological gap in the field of educational data mining. First, the paper provides a set of operational guidelines that guide researchers and educators from conception to model documentation, reducing reliance on purely intuitive approaches. Furthermore, the paper provides a critical analysis of the expert validation process and presents evidence of the technical challenges encountered during the modeling execution phase. Finally, the consolidated artifact serves as a decision-support instrument, enabling educational institutions to develop rigorous, reproducible early warning models tailored to the real needs of the academic context.Keywords:
Student dropout, Educational data mining, AI for Education.