DIGITAL LIBRARY
Q2A (QUESTION-TO-ANSWER): A MIXED-METHODS EVALUATION PIPELINE FOR AI-SUPPORTED EDUCATIONAL TOOLS: EVIDENCE FROM THE AI4EDU PROJECT
1 LuleƄ University of Technology (SWEDEN)
2 Institute for Language and Speech Processing, Athena Research Center (GREECE)
3 University of Cyprus (CYPRUS)
About this paper:
Appears in: EDULEARN26 Proceedings
Publication year: 2026
Article: 1347
ISBN: 978-84-09-88444-5
ISSN: 2340-1117
doi: 10.21125/edulearn.2026.1347
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Large Language Model (LLM)-based conversational assistants are increasingly entering classrooms, yet systematic evidence remains limited regarding the educational themes they effectively support and how users perceive and engage with them. The rapid deployment of artificial intelligence (AI) tools in education has introduced a critical methodological challenge: how can heterogeneous feedback data, such as interviews, Likert-scale questionnaires, open-ended survey responses, and system usage logs, be systematically integrated into a coherent, research-aligned evaluation process? While collecting feedback is common practice, structured methodologies for synthesizing diverse qualitative and quantitative evidence into collective conclusions remain underdeveloped. To address this gap, we introduce Q2A (Question-to-Answer), a three-stage mixed-methods evaluation pipeline developed within the AI4EDU project to assess two GenAI applications designed to support independent learning and instructional delivery: Study Buddy (student-facing) and Teacher Mate (teacher-facing) in upper secondary education. The Q2A pipeline consists of: (1) a human-centered, LLM-assisted mapping method aligning survey items with predefined research questions to ensure construct validity; (2) integrated qualitative and quantitative analyses, including thematic clustering of Likert-scale constructs, sentiment analysis (TextBlob, VADER, and BERT) of open-ended responses, and behavioral analytics derived from system logs; and (3) an aggregation stage that synthesizes findings across multiple analyses to produce unified, research-aligned conclusions. The Q2A pipeline was implemented across four European countries (Sweden, Greece, Cyprus, and Ireland), involving approximately 30 teachers and over 300 students. In the first stage, LLMs (GPT and Gemini) were used to generate mapping between specific questions from the survey instruments and the research questions designed by the project evaluators. These mappings were then reviewed and structured by human pedagogical experts, who also appended additional data sources to the mapping that should be considered for answering the research question(s). For example, evaluating the impact of Study Buddy on student motivation required aggregating survey responses, teacher feedback, usage logs, and academic performance indicators. The final aggregation stage revealed moderate to high positive effects on student motivation (3.7/5 in Sweden and Cyprus; 4.2/5 in Greece and Ireland). Importantly, results demonstrated that perceived educational value emerged from the convergence of survey data and behavioral indicators, highlighting how the mixed-methods aggregation of multiple feedback data sources provides greater interpretive clarity than isolated metrics alone. The Q2A pipeline offers a replicable and scalable framework for integrating multi-source educational data into systematic evaluation processes, providing practical guidance for researchers and policymakers assessing AI-supported learning environments.
Keywords:
Large Language Models (LLMs), Artificial Intelligence in Education (AIED), Conversational AI, GenAI, Sentiment Analysis, Thematic Analysis, Lexicon/Rule-Based Sentiment Analysis, Transformer-Based Sentiment Analysis, Educational Technology Assessment, AI4EDU.