USING LARGE LANGUAGE MODELS FOR DEVELOPING EDUCATIONAL ASSESSMENTS: AN EXPERIENCE REPORT OF AN IN-SERVICE TEACHER TRAINING PROGRAM IN GENERATIVE AI
Instituto Federal de Educação, Ciência e Tecnologia do Espírito Santo Campus Centro-Serrano (BRAZIL)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
The development of educational assessments is an essential yet burdensome task in teaching practice, especially given the high workload faced by teachers. In this context, recent advances in Generative Artificial Intelligence (GenAI), particularly Large Language Models (LLMs), have sparked international interest due to their potential to support pedagogical practices and reduce the operational demands of teaching activities. This study presents an experience report of an in-service teacher training program conducted in 2025, aimed at developing teachers' competencies in the use of LLMs for the creation of educational assessments. The training program involved 25 high school teachers from different subject areas and was structured into five sequential phases: Diagnosis, Training, Production, Application, and Evaluation. The Diagnosis phase aimed to understand how teachers create assessment items and their perceptions of GenAI in this process. Based on these findings, a three-hour in-person workshop was designed to introduce the use of LLMs for assessment creation. Driven by the interest shown during the workshop, the faculty chose to apply the acquired knowledge through the collective construction of a mock exam for the National High School Exam (ENEM), Brazil's primary higher education entrance examination. This exam was developed during the Production phase, where teachers used ChatGPT to create assessment items following the exam's style, employing One-shot and Prompt Engineering techniques. The exam was then conducted in the Application phase for 296 high school students. Subsequently, the Evaluation phase focused on the technological and pedagogical competencies acquired by the teachers. The Diagnosis and Evaluation phases provided valuable "before and after" insights. Initially, half of the teachers developed assessments using traditional individual strategies based on adapting existing materials, while the other half used AI tools occasionally but without pedagogical or technical methods, evidencing low literacy in these competencies. The group identified lack of time and the effort required to create high-quality, complex assessment items as their primary challenges. After the intervention, 75% of the teachers reported that AI significantly accelerated and facilitated the assessment creation process, although they emphasized the need for pedagogical review and additional contextualization of the generated items. Furthermore, following the training, all participants reported either beginning or increasing their use of AI tools in other pedagogical activities, such as lesson planning, instructional material development, and content study. These results suggest that structured training initiatives can support teachers in integrating Generative AI into assessment practices and contribute to the development of digital and pedagogical competencies. The study also highlights the potential for scalable replication of similar training programs in different educational contexts.Keywords:
Generative Artificial Intelligence, In-service Teacher Training, Learning Assessment, Large Language Models, Digital Literacy.