<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/CINECAstyle.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-20T10:37:00Z</responseDate><request verb="GetRecord" identifier="oai:iris.unitn.it:11572/458913" metadataPrefix="oai_dc">https://iris.unitn.it/oai/request</request><GetRecord><record><header><identifier>oai:iris.unitn.it:11572/458913</identifier><datestamp>2026-04-03T00:48:37Z</datestamp><setSpec>com_11572_237821</setSpec><setSpec>com_11572_101871</setSpec><setSpec>col_11572_237822</setSpec><setSpec>ou_ou00002</setSpec></header><metadata><oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:doc="http://www.lyncode.com/xoai" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
<dc:title>Learning without Labels - Reducing Supervision in Training, Inference, and Evaluation of Deep Neural Networks</dc:title>
<dc:creator>Conti, Alessandro</dc:creator>
<dc:contributor>Conti, Alessandro</dc:contributor>
<dc:contributor>Ricci, Elisa</dc:contributor>
<dc:contributor>Rota, Paolo</dc:contributor>
<dc:subject>Fine-tuning with domain shift, Inference without labels, Automatic benchmarking</dc:subject>
<dc:description>This thesis investigates how the reliance on supervision can be reduced across the entire deep learning pipeline. In the training phase, we explore unsupervised fine-tuning, focusing on Source-Free Unsupervised Domain Adaptation scenarios in visual tasks such as Facial Expression Recognition and video-based Action Recognition, primarily leveraging self-supervision and self-training. At inference, we address the challenge of removing fixed output vocabularies from Vision Language Models by formalizing the tasks of Vocabulary-free Image Classification and Vocabulary-free Semantic Segmentation and by introducing a family of efficient methods that adapt CLIP to the tasks. We also evaluate Large Multimodal Models under a similar constrained scenario, analyzing their predictions, categorizing their mistakes, and proposing tailored solutions to optimize their performance. Finally, we investigate unsupervised evaluation by proposing a framework that uses a Large Language Model and modular tools to automatically generate, execute, and interpret evaluation experiments for Large Multimodal Models without ground-truth labels. By reducing the need for human supervision at every stage of the deep learning pipeline, this thesis contributes toward a more flexible and efficient paradigm for developing and deploying deep neural networks in real-world, data-scarce, and open-ended settings.</dc:description>
<dc:date>2025-07-17</dc:date>
<dc:type>info:eu-repo/semantics/doctoralThesis</dc:type>
<dc:identifier>https://hdl.handle.net/11572/458913</dc:identifier>
<dc:identifier>http://dx.doi.org/10.15168/11572_458913</dc:identifier>
<dc:identifier>10.15168/11572_458913</dc:identifier>
<dc:language>eng</dc:language>
<dc:relation>firstpage:1</dc:relation>
<dc:relation>lastpage:195</dc:relation>
<dc:relation>numberofpages:195</dc:relation>
<dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
<dc:publisher>Università degli studi di Trento</dc:publisher>
<dc:publisher>place:TRENTO</dc:publisher>
<dc:rights>license:Creative commons</dc:rights>
<dc:rights>license uri:http://creativecommons.org/licenses/by-nc-nd/4.0/</dc:rights>
</oai_dc:dc></metadata></record></GetRecord></OAI-PMH>