The broad adoption of Large Language Models (LLMs) has transformed the landscape of Natural Language Processing (NLP) since the early 2020s, placing this PhD thesis within one of the fastest-growing areas of recent years. This research project first investigates the capabilities of language models for the NLP task of Information Extraction via prompting, and subsequently deepens our understanding of their internal mechanisms through neural probing. The first part of this PhD thesis advances the extraction of structured information from verbose textual documents through prompt engineering. It leverages different prompting techniques to extend the traditional, syntax-driven task of Information Extraction. In particular, the semantic understanding and prompting flexibility of LLMs are exploited to extract triples centred on Environmental, Social, and Governance aspects from companies’ non-financial reports. The resulting structured information is then represented as graphs, enabling descriptive, similarity-based, correlation, and interpretability analyses that uncover non-trivial insights into companies’ disclosures. Subsequently, the research questions of this PhD thesis move beyond the use of LLMs as black-box tools and instead focus on the limited understanding of their internal representations. Specifically, the first probing-oriented work introduces a novel methodology for investigating the knowledge-resolution process in LLMs by extracting factual structures from their latent space and visualizing their layer-wise dynamics within graph representations. Conversely, the final part of this thesis investigates the semantics encoded in the vector space of LLMs from a different perspective by addressing research questions related to the output-oriented interpretability of the semantic concepts represented within LLMs. It introduces a new probing paradigm, combining ideas from neural probing and symbolic representation, to generalise the extraction of structured knowledge from neural embeddings. By exploiting the properties of Vector Symbolic Architectures and hypervector algebra, this probing method provides non-trivial insights into the latent space of LLMs, ranging from concept-based patterns across input types to the investigation of representational changes before and after text generation. Following the research developments in the community, the novel methodologies proposed in this PhD thesis advance both the application of LLMs to traditional tasks and the semantic interpretability of their latent representations.
From Text to Neural Embeddings: Structuring Conceptual Knowledge in Large Language Models / Bronzini, M.. - (2026 Oct 02).
From Text to Neural Embeddings: Structuring Conceptual Knowledge in Large Language Models
Bronzini, Marco
2026-10-02
Abstract
The broad adoption of Large Language Models (LLMs) has transformed the landscape of Natural Language Processing (NLP) since the early 2020s, placing this PhD thesis within one of the fastest-growing areas of recent years. This research project first investigates the capabilities of language models for the NLP task of Information Extraction via prompting, and subsequently deepens our understanding of their internal mechanisms through neural probing. The first part of this PhD thesis advances the extraction of structured information from verbose textual documents through prompt engineering. It leverages different prompting techniques to extend the traditional, syntax-driven task of Information Extraction. In particular, the semantic understanding and prompting flexibility of LLMs are exploited to extract triples centred on Environmental, Social, and Governance aspects from companies’ non-financial reports. The resulting structured information is then represented as graphs, enabling descriptive, similarity-based, correlation, and interpretability analyses that uncover non-trivial insights into companies’ disclosures. Subsequently, the research questions of this PhD thesis move beyond the use of LLMs as black-box tools and instead focus on the limited understanding of their internal representations. Specifically, the first probing-oriented work introduces a novel methodology for investigating the knowledge-resolution process in LLMs by extracting factual structures from their latent space and visualizing their layer-wise dynamics within graph representations. Conversely, the final part of this thesis investigates the semantics encoded in the vector space of LLMs from a different perspective by addressing research questions related to the output-oriented interpretability of the semantic concepts represented within LLMs. It introduces a new probing paradigm, combining ideas from neural probing and symbolic representation, to generalise the extraction of structured knowledge from neural embeddings. By exploiting the properties of Vector Symbolic Architectures and hypervector algebra, this probing method provides non-trivial insights into the latent space of LLMs, ranging from concept-based patterns across input types to the investigation of representational changes before and after text generation. Following the research developments in the community, the novel methodologies proposed in this PhD thesis advance both the application of LLMs to traditional tasks and the semantic interpretability of their latent representations.I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



