EVALUATING LARGE LANGUAGE MODELS FOR EDUCATIONAL USE IN BEEKEEPING: A COMPARATIVE PERFORMANCE ANALYSIS
Agricultural University of Athens (GREECE)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
The rapid development of Artificial Intelligence, and particularly Large Language Models (LLMs), has significantly transformed the way information is accessed, processed, and used. In educational settings, LLMs can support learning, knowledge acquisition, and assessment by generating fast and context-sensitive responses. However, their performance may vary considerably across domains and task types, making systematic evaluation necessary.
This study examines the use of LLMs for educational purposes in beekeeping through a comparative evaluation of five models: LLaMA, Qwen, DeepSeek, MiniOrca, and Phi. The research specifically focused on assessing their effectiveness in understanding beekeeping content, answering educational questions, and generating valid and reliable subject-specific knowledge. To support this evaluation, the models were exposed to a corpus of thirteen beekeeping books. A structured testing framework was then developed, including multiple-choice, true/false, and open-ended questions based on university-level beekeeping course material.
The methodology was based on the collection and analysis of quantitative data regarding the accuracy and reliability of the responses generated by each model. The results revealed substantial differences in performance. LLaMA achieved the highest overall score, with an 85% success rate, followed by Phi (65%), DeepSeek (60%), Qwen (55%), and MiniOrca (25%). Analysis by question type showed that LLaMA performed particularly well in true/false questions, whereas the remaining models demonstrated greater instability and more limitations, especially in open-ended responses.
The findings indicate that, despite the rapid progress of LLMs, important challenges remain in their educational use within specialized domains such as beekeeping, particularly in relation to knowledge generalization, accuracy, and coherence. This research contributes to the evaluation of LLMs for educational purposes in beekeeping and highlights the importance of careful model selection, as well as the potential value of combining multiple AI tools to provide high-quality and reliable educational support.
Acknowledgement:
This publication is part of the TALLHEDA project that has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No. 101136578. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them.Keywords:
Large Language Models, Artificial Intelligence, LLM evaluation, beekeeping education, educational technology, LLaMA, Phi, DeepSeek, Qwen, MiniOrca.