Sentiment Analysis (SA) models harbor inherent social biases that can be harmful in real-world applications. These biases are identified by examining the output of SA models for sentences that only vary in the identity groups of the subjects. Constructing natural, linguistically rich, relevant, and diverse sets of sentences that provide sufficient coverage over the domain is expensive, especially when addressing a wide range of biases: it requires domain experts and/or crowd-sourcing. In this paper, we present a novel bias testing framework, BTC-SAM, which generates high-quality test cases for bias testing in SA models with minimal specification using Large Language Models (LLMs) for the controllable generation of test sentences. Our experiments show that relying on LLMs can provide high linguistic variation and diversity in the test sentences, thereby offering better test coverage compared to base prompting methods even for previously unseen biases.

BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models / Kardkovács, Z.T., Djennane, L., Field, A., Benatallah, B., Gaci, Y., Casati, F., Gaaloul, W.. - (2025), pp. 15108-15124. (30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 chn 2025) [10.18653/v1/2025.emnlp-main.763].

BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models

Benatallah, Boualem;Casati, Fabio;
2025-01-01

Abstract

Sentiment Analysis (SA) models harbor inherent social biases that can be harmful in real-world applications. These biases are identified by examining the output of SA models for sentences that only vary in the identity groups of the subjects. Constructing natural, linguistically rich, relevant, and diverse sets of sentences that provide sufficient coverage over the domain is expensive, especially when addressing a wide range of biases: it requires domain experts and/or crowd-sourcing. In this paper, we present a novel bias testing framework, BTC-SAM, which generates high-quality test cases for bias testing in SA models with minimal specification using Large Language Models (LLMs) for the controllable generation of test sentences. Our experiments show that relying on LLMs can provide high linguistic variation and diversity in the test sentences, thereby offering better test coverage compared to base prompting methods even for previously unseen biases.
2025
EMNLP 2025 - 2025 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference
Suzhou, China
Association for Computational Linguistics (ACL)
Kardkovács, Zsolt T.; Djennane, Lynda; Field, Anna; Benatallah, Boualem; Gaci, Yacine; Casati, Fabio; Gaaloul, Walid
BTC-SAM: Leveraging LLMs for Generation of Bias Test Cases for Sentiment Analysis Models / Kardkovács, Z.T., Djennane, L., Field, A., Benatallah, B., Gaci, Y., Casati, F., Gaaloul, W.. - (2025), pp. 15108-15124. (30th Conference on Empirical Methods in Natural Language Processing, EMNLP 2025 chn 2025) [10.18653/v1/2025.emnlp-main.763].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/504650
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 2
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact