RELIABILITY OF ARTIFICIAL INTELLIGENCE AS AN EVALUATION TOOL IN ENGINEERING EDUCATION: A CASE STUDY IN COMPOSITE MATERIALS
1 Instituto Universitario de Tecnología de Materiales (IUTM), Universitat Politècnica de València (SPAIN)
2 Instituto de Diseño y Fabricación (IDF), Universitat Politècnica de València (SPAIN)
3 Universitat Politècnica de València (SPAIN)
4 Centro de Investigación Gestión e Ingeniería de Producción (CIGIP), Universitat Politècnica de València (SPAIN)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
Evaluating comprehensive engineering assignments is traditionally a time-intensive process. While Artificial Intelligence (AI) offers a potential solution for automated grading, its reliability compared to expert human evaluation requires rigorous validation. This study investigates the effectiveness of generative AI models in evaluating complex technical reports in higher education.
The study was conducted at the Alcoy Campus of the Universitat Politècnica de València (UPV) during the 2024/2025 academic year, involving 15 Mechanical Engineering students. Participants were tasked with researching the manufacturing processes of a commercial polymer matrix composite product. Their reports required a scientific and commercial analysis encompassing cost, quality, environmental impact, and recent technological advancements. To test evaluation reliability, an identical grading rubric was applied across three distinct assessment modalities: student peer review, expert instructor evaluation (serving as the control baseline), and automated AI assessment using two distinct platforms (Gemini and ChatGPT).
A comparative analysis of the final scores revealed an 85% agreement rate between the AI-generated evaluations and those provided by the course instructor.
The findings demonstrate that Large Language Models, when constrained and guided by well-defined rubrics, possess a remarkably high degree of reliability in assessing technical engineering coursework. This suggests that AI can serve as a highly effective, time-saving tool for educators, reliably augmenting or substituting traditional grading processes without compromising academic standards.Keywords:
Artificial Intelligence in Education, Automated Assessment, Peer Review, Engineering Education, Evaluation Rubrics.