DIGITAL LIBRARY
EXPLAINING CROSS-LINGUAL KNOWLEDGE TRANSFER FOR FAIR AUTOMATED ESSAY SCORING IN MULTILINGUAL LANGUAGE LEARNING
University of Alberta (CANADA)
About this paper:
Appears in: EDULEARN26 Proceedings
Publication year: 2026
Article: 2665 (abstract only)
ISBN: 978-84-09-88444-5
ISSN: 2340-1117
doi: 10.21125/edulearn.2026.2665
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Automated essay scoring (AES) is increasingly used to support large-scale, consistent, and cost-effective assessment, yet its performance across languages and scoring criteria remains insufficiently understood. This challenge is particularly important in multilingual education, where annotated learner corpora are unevenly distributed and lower-resource languages may be underserved. This study examines how effectively multilingual and language-agnostic pre-trained language models transfer AES knowledge across German, Italian, and Czech to support fairer assessment of second-language writing. Specifically, it asks whether cross-lingual transfer occurs uniformly across scoring dimensions such as grammatical accuracy, orthography, vocabulary control, coherence, and sociolinguistic appropriateness.

The study uses essays from the MERLIN (Multilingual Platform for European Reference Levels: Interlanguage Exploration in Context) corpus, scored by human raters using the Common European Framework of Reference (CEFR). Language-agnostic sentence-embedding models and multilingual transformer models are fine-tuned on essays in one language and tested on essays in another. Performance is examined at both the holistic and rubric levels using standard classification metrics. To make model behaviour more interpretable, the study combines quantitative evaluation with neural representation analysis using t-distributed stochastic neighbor embedding (t-SNE) and qualitative error analysis. This design helps identify which writing features transfer effectively across languages, where model bias may emerge, and how representation spaces differ across model families.

Preliminary analyses suggest that language-agnostic models support stronger cross-lingual transfer for higher-level criteria such as coherence, while multilingual transformer models remain more competitive on language-dependent features such as grammar and orthography. Early visualization results also indicate that language-agnostic models produce more overlapping cross-language representation spaces, which may support better generalization when training data are limited.

The study contributes empirical evidence and practical guidance for the design of multilingual AES systems in educational settings. Its implications are relevant for language learning innovation, online assessment, and the responsible use of artificial intelligence in education, particularly where fair assessment must be achieved with limited annotated data.
Keywords:
Automated Essay Scoring, Multilingual Education, Language Assessment, Artificial Intelligence in Education, Student Learning Assessment, Fairness, Language Learning.