We conduct a systematic audit of three widely used social reasoning benchmarks, SocialIQa, FauxPas-EAI, and ToMi, and uncover pervasive flaws in both benchmark items and evaluation methodology. Using five LLMs (GPT-3, 3.5, 4, o1, and LLaMA 3.1) as diagnostic tools, we identify structural, semantic, and pragmatic issues in benchmark design (e.g., duplicated items, ambiguous wording, and implausible answers), as well as scoring procedures that prioritize output form over the reasoning process. Through systematic human annotation and reevaluation on cleaned benchmark subsets, we find that model scores often improve not due to due to erratic surface wording variations and not to improved reasoning. In fact, further analyses show that model performance is highly sensitive to minor input variations such as context availability and phrasing, revealing that high scores may reflect alignment with format-specific cues rather than consistent inference based on the input. These findings challenge the validity of current benchmark-based claims about social reasoning in LLMs, and highlight the need for evaluation protocols that assess reasoning as a process of drawing inference from available information, rather than as static output selection. We release audited data and evaluation tools to support more interpretable

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It / Mousavi, S.M., Cecchinato, E., Horníková, L., Riccardi, G.. - (2026), pp. 1747-1759. (19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 MOROCCO march 2026) [10.18653/v1/2026.findings-eacl.89].

Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It

Mousavi, Seyed Mahed
Primo
;
Riccardi, Giuseppe
2026-01-01

Abstract

We conduct a systematic audit of three widely used social reasoning benchmarks, SocialIQa, FauxPas-EAI, and ToMi, and uncover pervasive flaws in both benchmark items and evaluation methodology. Using five LLMs (GPT-3, 3.5, 4, o1, and LLaMA 3.1) as diagnostic tools, we identify structural, semantic, and pragmatic issues in benchmark design (e.g., duplicated items, ambiguous wording, and implausible answers), as well as scoring procedures that prioritize output form over the reasoning process. Through systematic human annotation and reevaluation on cleaned benchmark subsets, we find that model scores often improve not due to due to erratic surface wording variations and not to improved reasoning. In fact, further analyses show that model performance is highly sensitive to minor input variations such as context availability and phrasing, revealing that high scores may reflect alignment with format-specific cues rather than consistent inference based on the input. These findings challenge the validity of current benchmark-based claims about social reasoning in LLMs, and highlight the need for evaluation protocols that assess reasoning as a process of drawing inference from available information, rather than as static output selection. We release audited data and evaluation tools to support more interpretable
2026
Findings of the Association for Computational Linguistics: EACL 2026
Morocco
Association for Computational Linguistics (ACL)
9798891763869
Mousavi, Seyed Mahed; Cecchinato, Edoardo; Horníková, Lucia; Riccardi, Giuseppe
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It / Mousavi, S.M., Cecchinato, E., Horníková, L., Riccardi, G.. - (2026), pp. 1747-1759. (19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 MOROCCO march 2026) [10.18653/v1/2026.findings-eacl.89].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/485291
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex 0
social impact