AI chatbots are increasingly used by students as study tools in physics, raising practical questions about their reliability on conceptual tasks. Existing evaluations of large language models (LLMs) on physics concept inventories rely almost exclusively on instruments that have been publicly available for years and likely appear in model training data, making it difficult to disentangle physics competence from familiarity with the test items themselves. We address this by evaluating three LLMs (GPT-5.2, Gemini 3 Pro, Gemini 3 Flash) on the Classical Relativity Concept Inventory, a recently developed and validated 21-item instrument on Galilean relativity that was not publicly available at the time of testing. Each item was administered 30 times per model, and all 1890 responses were qualitatively coded along three dimensions: visual interpretation, physics reasoning, and coordination. Because GPT-5.2 was tested in a no-reasoning configuration whereas the Gemini models were tested with their provider-default reasoning settings, cross-model accuracy comparisons are treated descriptively rather than as direct evidence of differential physics competence. Mean accuracy was 97% for Gemini 3 Flash, 89% for Gemini 3 Pro, and 73% for GPT-5.2 in the no-reasoning configuration (rising to 85%–86% when reasoning was enabled), compared to 62% for the student sample (N = 267). However, on a small number of items GPT-5.2 and Gemini 3 Pro drop to near-zero accuracy, while Gemini 3 Flash, although more robust, still misreads specific diagrams. The qualitative analysis shows that these failures stem predominantly from misinterpretations of visual content rather than from deficits in physics knowledge, and that LLM errors differ structurally from those of students: when models err, they converge on a single distractor with high consistency, whereas student errors are more broadly distributed. These findings indicate that chatbot reliability on conceptual physics is item-dependent and unpredictable, with direct implications for how concept inventories are administered.

Performance and failure modes of AI chatbots on a novel concept inventory on relativity in classical mechanics / Tufino, E., Giovanzana, C., Zamboni, A., Onorato, P., Oss, S.. - In: EUROPEAN JOURNAL OF PHYSICS. - ISSN 0143-0807. - 47:4(2026), p. 045704. [10.1088/1361-6404/ae8c90]

Performance and failure modes of AI chatbots on a novel concept inventory on relativity in classical mechanics

Tufino, Eugenio;Giovanzana, Caterina;Zamboni, Andrea;Onorato, Pasquale;Oss, Stefano
2026-01-01

Abstract

AI chatbots are increasingly used by students as study tools in physics, raising practical questions about their reliability on conceptual tasks. Existing evaluations of large language models (LLMs) on physics concept inventories rely almost exclusively on instruments that have been publicly available for years and likely appear in model training data, making it difficult to disentangle physics competence from familiarity with the test items themselves. We address this by evaluating three LLMs (GPT-5.2, Gemini 3 Pro, Gemini 3 Flash) on the Classical Relativity Concept Inventory, a recently developed and validated 21-item instrument on Galilean relativity that was not publicly available at the time of testing. Each item was administered 30 times per model, and all 1890 responses were qualitatively coded along three dimensions: visual interpretation, physics reasoning, and coordination. Because GPT-5.2 was tested in a no-reasoning configuration whereas the Gemini models were tested with their provider-default reasoning settings, cross-model accuracy comparisons are treated descriptively rather than as direct evidence of differential physics competence. Mean accuracy was 97% for Gemini 3 Flash, 89% for Gemini 3 Pro, and 73% for GPT-5.2 in the no-reasoning configuration (rising to 85%–86% when reasoning was enabled), compared to 62% for the student sample (N = 267). However, on a small number of items GPT-5.2 and Gemini 3 Pro drop to near-zero accuracy, while Gemini 3 Flash, although more robust, still misreads specific diagrams. The qualitative analysis shows that these failures stem predominantly from misinterpretations of visual content rather than from deficits in physics knowledge, and that LLM errors differ structurally from those of students: when models err, they converge on a single distractor with high consistency, whereas student errors are more broadly distributed. These findings indicate that chatbot reliability on conceptual physics is item-dependent and unpredictable, with direct implications for how concept inventories are administered.
2026
4
Tufino, Eugenio; Giovanzana, Caterina; Zamboni, Andrea; Onorato, Pasquale; Oss, Stefano
Performance and failure modes of AI chatbots on a novel concept inventory on relativity in classical mechanics / Tufino, E., Giovanzana, C., Zamboni, A., Onorato, P., Oss, S.. - In: EUROPEAN JOURNAL OF PHYSICS. - ISSN 0143-0807. - 47:4(2026), p. 045704. [10.1088/1361-6404/ae8c90]
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/497191
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact