DIGITAL LIBRARY
A QUALITY ASSURANCE FRAMEWORK FOR AI-SUPPORTED CONSTRUCTION OF INTELLECTUAL-SKILL TEST ITEMS
1 The Open University of Japan (JAPAN)
2 Minna Learning Environment Laboratory (JAPAN)
3 Coach and Research Inc. (JAPAN)
4 Arbege Corporation (JAPAN)
About this paper:
Appears in: EDULEARN26 Proceedings
Publication year: 2026
Article: 1602
ISBN: 978-84-09-88444-5
ISSN: 2340-1117
doi: 10.21125/edulearn.2026.1602
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Improving the quality of online assessment remains a major challenge in higher education, where many examinations still place too much emphasis on rote memorization. An empirical analysis of 7,517 test items used at a Japanese distance-learning university showed that about 79% of the items targeted simple knowledge recall, while items assessing intellectual skills were relatively scarce. One major reason is that intellectual-skill items are difficult to construct: they require learners to apply rules or principles in new contexts, yet many instructors lack explicit instructional design (ID) knowledge.

To address this problem, the present research consists of three interrelated components for AI-supported construction of intellectual-skill test items. The first component examines how AI can be guided to extract learning objectives from textbook data and classify them using Robert Gagné's learning-outcome categories. The second component examines how AI can be guided to generate new, analogous, and improved items aligned with those objectives. The third component investigates a quality assurance mechanism in which an Auditor AI supervises a Worker AI, evaluates generated items, and provides revision feedback when needed. In this presentation, we report a partial implementation of the third component, building on the procedures and checklist system developed in the second component.

The second component produced a structured checklist system for evaluating generated items. It consists of a core checklist for general item-writing principles and sub-checklists tailored to learning-outcome types and three item levels: knowledge recall, lower-order intellectual skills, and higher-order intellectual skills. The criteria cover objective alignment, match with the specified item level, sufficient problem-solving context, and avoidance of superficial cues.

The third component examines how an Auditor AI uses this checklist system to guide a Worker AI in generating and revising items. The Worker AI received specifications and checklist-based guidance to produce items, while the Auditor AI reviewed the drafts against the same criteria and indicated how they should be revised. Both roles were instantiated using existing AI systems and operated by the researchers under experimental conditions, rather than as autonomous agents.

GPT, Gemini, and Claude were used as Worker AIs to leverage their differing generative characteristics. To examine cross-model applicability, the same materials, specifications, and procedures were applied across the three LLMs through generation–audit cycles. The item drafts produced by the three models showed clear differences: some produced elaborate scenarios whereas others adhered more closely to specifications, and some drifted toward knowledge-recall items whereas others made stronger attempts to construct intellectual-skill items. The Auditor AI evaluated these drafts against the checklist criteria and identified issues such as insufficient context for problem solving, weak objective alignment, and item-level mismatches, and indicated how they should be revised. Further work is needed on iterative revision and stable production of finalized items. These findings suggest that the proposed approach can serve as a practical support tool for faculty and as a basis for quality assurance across different LLMs, helping shift item construction from knowledge-recall questions toward items that better assess intellectual skills.
Keywords:
Instructional Design, Assessment Design, Intellectual Skills, AI-Supported Item Construction.