AGENTIC AI FOR RUBRIC-ALIGNED FEEDBACK ON HANDWRITTEN STUDENT ESSAYS: A MIXED-METHODS EVALUATION
1 Singapore University of Technology and Design (SUTD) (SINGAPORE)
2 University of Information Technology (UIT) (MYANMAR)
3 The Chinese University of Hong Kong (Shenzhen) (CHINA)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Providing detailed feedback on student writing is a key component of effective teacher–student learning interactions, yet generating such feedback remains one of the most time-consuming tasks for educators. Recent advances in large language models (LLMs) have enabled automated writing feedback at scale, but most existing approaches rely on single-pass prompting, where the model evaluates an essay and generates feedback in a single step. While agentic AI systems have shown promise on complex reasoning tasks, their effectiveness for formative writing feedback remains underexplored.
This study investigates whether an agentic artificial intelligence framework can improve rubric alignment and pedagogical usefulness in automated feedback for handwritten student essays. Specifically, the study addresses two research questions:
(1) Can an agentic AI framework improve rubric alignment in automated essay feedback? and
(2) Does decomposing essay evaluation into structured stages produce more pedagogically actionable feedback than single-pass LLM prompting?
The proposed system, TinyMark, developed by Tiny Equations, decomposes essay evaluation into multiple structured stages designed to mimic a teacher’s cognitive workflow, including rubric-based analysis, issue identification, and feedback generation. The evaluation used a dataset of twenty handwritten essays from an undergraduate academic writing course where English was used as a second language. The study was conducted in two phases: a pilot evaluation of five essays used to shortlist candidate systems and a main evaluation of fifteen essays comparing TinyMark with GPT-5.4 Thinking, Gemini 3 Thinking, and Claude Opus 4.5. All systems received identical essay images, rubric criteria, and task instructions.
A mixed-methods evaluation framework combining quantitative scoring and qualitative analysis was employed. Human raters assessed rubric alignment across five writing dimensions and feedback quality across specificity, actionability, accuracy, depth, and pedagogical tone. A secondary LLM-as-judge analysis examined the effect of evaluator identity on ranking outcomes.
Results from the main human evaluation did not support the claim that the agentic framework improved rubric alignment. TinyMark scored lower on rubric alignment than the three frontier baseline models, particularly on argument development and academic style. However, differences in feedback quality were considerably smaller. TinyMark performed comparably on accuracy, tone, and specificity, suggesting that rubric coverage alone may not fully capture the pedagogical value of feedback. The findings highlight both the limitations and potential of agentic decomposition for formative writing assessment and underscore the importance of human evaluation when comparing AI-generated feedback systems.Keywords:
Artificial intelligence in education, automated essay feedback, agentic AI, large language models, formative assessment, writing assessment, teacher workload reduction.