DIGITAL LIBRARY
ASSESSMENT ARTIFACTS IN STEM FOR THE AGE OF GENERATIVE AI: MINIMUM VIABLE EVIDENCE, VALIDATION METACOGNITION, AND DISCIPLINARY JUDGMENT
National University (UNITED STATES)
About this paper:
Appears in: EDULEARN26 Proceedings
Publication year: 2026
Article: 1985
ISBN: 978-84-09-88444-5
ISSN: 2340-1117
doi: 10.21125/edulearn.2026.1985
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
We examine the design of assessment artifacts in STEM that remain meaningful in the age of Generative AI. As AI systems increasingly produce polished text, code, explanations, and solutions to problems, traditional product-based assignments are becoming less reliable indicators of student understanding. The challenge is especially acute in STEM, where correctness alone may mask shallow reasoning, unverified assumptions, or unexamined dependence on machine-generated outputs. In response, this paper argues for a shift from evaluating final answers toward evaluating disciplinary judgment, validation practices, and process-based evidence of learning.

The article proposes assessment artifacts structured to make student thinking visible and to foreground supervision rather than mere production. These include flawed-proof critiques, annotated revisions of AI-generated solutions, debugging rationales, parameter sensitivity analyses, lab decision logs, oral defense components, uncertainty interpretation tasks, and AI audit trails documenting prompts, outputs, corrections, and justification. Such artifacts are not designed to ban AI, but to require students to engage critically with it by identifying errors, validating claims, explaining revisions, and demonstrating conceptual ownership of the work. The paper extends this argument through the idea of Minimum Viable Evidence (MVE), asking what the smallest but sufficient body of evidence is to show that a student understands, can verify, and can defend the submitted work within a disciplinary context.

We focus particularly on STEM because these disciplines depend heavily on traceable reasoning, methodological rigor, and explicit relationships among assumptions, procedures, and results. In these settings, a student’s ability to recognize when an output is plausible but wrong is often more educationally significant than the output itself. Assessment artifacts that require evaluation, correction, and reflective explanation, therefore, offer a more robust way to measure learning than assignments completed through fluent AI generation.

Rather than treating STEM as a single category, the paper examines how MVE must be adapted across disciplines. In mathematics, evidence may center on proof structure, justified steps, and recognition of invalid inferences. In engineering, it may lie in tradeoff reasoning, constraint awareness, and the defense of design decisions. In computing, it may be shown through testing strategy, debugging logic, revision history, and explanation of why code was changed rather than whether it runs. In the experimental sciences, it may depend on the coherence among hypotheses, methods, data quality, uncertainty, and interpretation. The paper also argues for a deeper analysis of submissions, focusing on coherence across sections, alignment between the method and the conclusion, and similarity and lexicographical scores across submissions indications of replicated reasoning.

Rather than presenting AI as an external threat, we treat it as a catalyst for assessment redesign. The goal is to preserve rigor while aligning evaluation with the realities of AI-mediated work in contemporary academic and professional environments. We conclude that well-designed STEM assessment artifacts can help faculty move beyond AI policing toward more authentic and cognitively demanding forms of assessment that capture how students think, verify, and take responsibility for machine-assisted work.
Keywords:
Assessments, AI-proof Assessments, Agentic AI, Generative AI, Education in Transition.