Linear properties are ubiquitous in the representations of language models; however, testing them experimentally remains a challenging task. This work focuses on relational linearity: the hypothesis that, for a fixed relation (e.g., “plays”), the unembedding of an object (e.g., “trumpet”) can be predicted from the embedding of its subject (e.g., “Miles Davis”) by a linear map. We present an experimental method to test the formulation of relational linearity by Marconato et al. (2024). Specifically, we introduce a probing method, based on Kullback-Leibler divergence, to evaluate this property and examine its variation across layers and paraphrased relational queries. It is also more efficient than previous work; for example, it avoids the crude Jacobian approximations used in Linear Relational Embeddings by Hernandez et al. (2024). Our findings across four datasets show that relational linearity varies across models, exhibits layer-wise patterns consistent with prior observations about linguistic information in model representations, and is differently affected by changes in how the relation is phrased.

Relational Linear Properties in Language Models: An Empirical Investigation / Valer, G., Gresele, L., Bronzini, M., Marconato, E.. - ELETTRONICO. - (2026). (International Conference on Machine Learning Seoul, South Korea 11/07/2026) [10.5281/zenodo.21992706].

Relational Linear Properties in Language Models: An Empirical Investigation

Valer, Giovanni
Primo
;
Gresele, Luigi
Co-ultimo
;
Bronzini, Marco
Co-ultimo
;
Marconato, Emanuele
Co-ultimo
2026-01-01

Abstract

Linear properties are ubiquitous in the representations of language models; however, testing them experimentally remains a challenging task. This work focuses on relational linearity: the hypothesis that, for a fixed relation (e.g., “plays”), the unembedding of an object (e.g., “trumpet”) can be predicted from the embedding of its subject (e.g., “Miles Davis”) by a linear map. We present an experimental method to test the formulation of relational linearity by Marconato et al. (2024). Specifically, we introduce a probing method, based on Kullback-Leibler divergence, to evaluate this property and examine its variation across layers and paraphrased relational queries. It is also more efficient than previous work; for example, it avoids the crude Jacobian approximations used in Linear Relational Embeddings by Hernandez et al. (2024). Our findings across four datasets show that relational linearity varies across models, exhibits layer-wise patterns consistent with prior observations about linguistic information in model representations, and is differently affected by changes in how the relation is phrased.
2026
Mechanistic Interpretability Workshop at ICML 2026
Seoul, South Korea
International Conference on Machine Learning
Valer, Giovanni; Gresele, Luigi; Bronzini, Marco; Marconato, Emanuele
Relational Linear Properties in Language Models: An Empirical Investigation / Valer, G., Gresele, L., Bronzini, M., Marconato, E.. - ELETTRONICO. - (2026). (International Conference on Machine Learning Seoul, South Korea 11/07/2026) [10.5281/zenodo.21992706].
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/506050
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact