BIAS, CREDIBILITY, AND RELIABILITY IN MULTILINGUAL CHATGPT PROMPTING
1 South Mediterranean University (TUNISIA)
2 Mediterranean Institute of Technology (TUNISIA)
About this paper:
Conference name: 18th International Conference on Education and New Learning Technologies
Dates: 29 June-1 July, 2026
Location: Palma, Spain
Abstract:
As digital technologies and artificial intelligence (AI) become increasingly embedded in educational systems worldwide, their influence on teaching and learning continues to expand. AI-powered tools such as large language models are now used to support instruction, generate educational content and facilitate multilingual communication. However, alongside their growing adoption in global education contexts, concerns have emerged regarding bias, credibility, reliability, and cultural representation, particularly in linguistically and culturally diverse environments. These concerns are especially significant within the broader framework of global issues in education, where access to trustworthy knowledge remain central priority. This study situates the use of AI chatbots within the global education context, focusing specifically on the performance of ChatGPT 4.o across multiple languages and culturally sensitive topics. As educational institutions increasingly rely on AI tools to support students, understanding how these systems function across linguistic and cultural contexts becomes essential. The research investigates potential bias in ChatGPT’s responses to multilingual question prompts addressing culturally sensitive and historical issues. Such topics are particularly relevant in education, therefore, evaluating how AI systems handle these topics in different languages is critical for assessing their appropriateness as educational tools. To conduct this analysis, a set of prompts was initially developed in English and subsequently translated into Spanish, Arabic, French, and German. These languages were selected to represent diverse linguistic families and geopolitical contexts. By comparing responses generated in each language, the study assesses language structures, themes, and data presentation. Similarity rates were applied to measure the degree of alignment across responses, providing a quantitative basis for evaluating cross-linguistic reliability. The findings indicate that ChatGPT 4.o generally produces highly consistent responses across languages, with similarity rates exceeding 0.85 in several cases. However, the study also identifies variations in similarity rates, reflecting subtle differences in how certain topics are interpreted and articulated across languages. Moreover, the study emphasizes that bias in AI is not always overt or explicit; it may manifest subtly. Ensuring that AI tools align with principles of accuracy and balanced representation is therefore essential for their responsible integration into curricula worldwide. While ChatGPT 4.o demonstrates a high degree of cross-linguistic consistency, the observed variations underscore the importance of critically assessing AI-generated content in global education settings. Addressing concerns related to bias, credibility, and reliability is essential for maximizing the educational benefits of AI technologies while minimizing risks.Keywords:
Artificial Intelligence, Bias, ChatGPT, Linguistics Contexts, Prompt Engineering.