<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/CINECAstyle.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-20T13:35:04Z</responseDate><request verb="GetRecord" identifier="oai:iris.unitn.it:11572/483952" metadataPrefix="oai_dc">https://iris.unitn.it/oai/request</request><GetRecord><record><header><identifier>oai:iris.unitn.it:11572/483952</identifier><datestamp>2026-05-25T15:54:04Z</datestamp><setSpec>com_11572_237821</setSpec><setSpec>com_11572_101871</setSpec><setSpec>col_11572_237822</setSpec><setSpec>ou_ou00002</setSpec></header><metadata><oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:doc="http://www.lyncode.com/xoai" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
<dc:title>Models and Application of Question Retrieval for Natural Language Processing</dc:title>
<dc:creator>Campese, Stefano</dc:creator>
<dc:contributor>Campese, Stefano</dc:contributor>
<dc:contributor>Moschitti, Alessandro</dc:contributor>
<dc:subject>Question Retrieval, Semantic Equivalence, Database Question Answering, Question Ranking, Self-Supervised Pre-training, LLM Coherence, Retrieval-Augmented Generation, Retrieval Coherence, Dense Retrieval, Dataset Declassification, Privacy-Preserving NLP, Answer Sentence Selection</dc:subject>
<dc:description>This thesis investigates the role of question understanding in Question Answering systems, developing methods that exploit question semantic equivalence at progressively larger scales: from individual question pairs, through equivalence clusters, to entire datasets. The first part addresses question retrieval at scale. We introduce QUADRo, a retrieval framework operating over millions of question-answer pairs, and the Question Ranking Corpus (QRC), a large-scale resource with answer-aware annotations and challenging hard negatives. We demonstrate that incorporating answers during retrieval substantially improves accuracy, as answers serve as a semantic bridge between questions that share little lexical overlap but seek the same information. To reduce annotation costs, we develop Question Ranking Pre-training (QRP), a self-supervised method that learns question equivalence patterns without labeled data, achieving significant improvements while reducing model variance by over 50\%. The second part extends pairwise equivalence to question clusters. We analyze coherence in Large Language Models, finding that a substantial portion of question clusters exhibit incoherent behavior: models answer some phrasings correctly while failing on semantically equivalent alternatives. This reveals that understanding failures, not just knowledge gaps, limit LLM performance. We introduce Question-Augmented Generation (q-RAG), which supplements prompts with retrieved similar questions, improving accuracy by up to 9 percentage points and coherence by up to 28 points. We further show that q-RAG's benefits can be distilled into model parameters through Direct Preference Optimization (DPO) and Supervised Fine-Tuning, producing standalone models with improved coherence that surpass the inference-time approach. For retrieval systems, we apply clusters to train models for consistency: the Coherence Ranking Loss improves ranking coherence by up to 30\% while simultaneously improving relevance. The third part lifts equivalence to the dataset level. We introduce dataset declassification, a framework that replaces proprietary questions with semantically equivalent public alternatives, enabling dataset sharing without exposing sensitive content. Models trained on fully declassified data match baseline performance (WikiQA $\Delta \approx 0$, TrecQA $|\Delta| \leq 1.2$  points), and test set declassification preserves evaluation validity when high-quality mappings exist ($|\Delta| \leq 2$ on standard benchmarks),  enabling the release of ``shadow benchmarks'' for evaluation integrity. We identify boundary conditions through experiments on adversarially-constructed benchmarks. Together, these contributions show that question semantic equivalence, systematically exploited at multiple scales, enables substantial improvements to QA system accuracy, consistency, and evaluation integrity.</dc:description>
<dc:date>2026-04-27</dc:date>
<dc:type>info:eu-repo/semantics/doctoralThesis</dc:type>
<dc:identifier>https://hdl.handle.net/11572/483952</dc:identifier>
<dc:language>eng</dc:language>
<dc:relation>firstpage:1</dc:relation>
<dc:relation>lastpage:203</dc:relation>
<dc:relation>numberofpages:203</dc:relation>
<dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
<dc:publisher>Università degli studi di Trento</dc:publisher>
<dc:publisher>place:TRENTO</dc:publisher>
<dc:rights>license:Tutti i diritti riservati (All rights reserved)</dc:rights>
<dc:rights>license uri:iris.PRI01</dc:rights>
</oai_dc:dc></metadata></record></GetRecord></OAI-PMH>