<?xml version="1.0" encoding="UTF-8"?><?xml-stylesheet type="text/xsl" href="static/CINECAstyle.xsl"?><OAI-PMH xmlns="http://www.openarchives.org/OAI/2.0/" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/ http://www.openarchives.org/OAI/2.0/OAI-PMH.xsd"><responseDate>2026-09-21T22:36:41Z</responseDate><request verb="GetRecord" identifier="oai:iris.unitn.it:11572/484810" metadataPrefix="oai_dc">https://iris.unitn.it/oai/request</request><GetRecord><record><header><identifier>oai:iris.unitn.it:11572/484810</identifier><datestamp>2026-06-12T09:45:51Z</datestamp><setSpec>com_11572_237821</setSpec><setSpec>com_11572_101871</setSpec><setSpec>col_11572_237822</setSpec><setSpec>ou_ou00002</setSpec></header><metadata><oai_dc:dc xmlns:oai_dc="http://www.openarchives.org/OAI/2.0/oai_dc/" xmlns:doc="http://www.lyncode.com/xoai" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xmlns:dc="http://purl.org/dc/elements/1.1/" xsi:schemaLocation="http://www.openarchives.org/OAI/2.0/oai_dc/ http://www.openarchives.org/OAI/2.0/oai_dc.xsd">
<dc:title>Language-Grounded Post-Completion Mistake Detection in Procedural Videos</dc:title>
<dc:creator>Loginova, Olga</dc:creator>
<dc:contributor>Loginova, Olga</dc:contributor>
<dc:contributor>Passerini, Andrea</dc:contributor>
<dc:contributor>Ricci, Elisa</dc:contributor>
<dc:contributor>Staiano, Jacopo</dc:contributor>
<dc:description>Mistake detection in procedural videos is the task of identifying errors in activities such as cooking, assembly, or repair. The domain represents a critical yet underexplored challenge. This thesis focuses on Post-Completion Mistake Detection (PCMD), where a model must verify a full procedure execution and localize deviations from the intended protocol. PCMD is under-researched and still held back by fragmented error taxonomies, staged and scarce datasets, and complex, computationally demanding, often domain-specific vision-first models. &#xd;
This thesis develops a unified, language-centered PCMD framework. First, it establishes the limitations of end-to-end Vision-Language Models (VLMs) for procedural verification. Through gaps in temporal reasoning of ongoing and completed actions, failures in understanding of cause-effect relations in procedural structures, and model tendencies towards ``blind guessing'', the thesis demonstrates that VLMs struggle with fine-grained temporal logic. The diagnostics prove that reliable mistake detection requires structured and interpretable mechanisms over black-box VLM reasoning alone. Second, to address the data bottleneck, the thesis introduces PIE-V, a semi-synthetic pipeline for generating mistake-aware datasets. Using psychology-informed error planning, PIE-V injects semantic mistakes into clean procedures. It delivers controllable, error-rich variants that approximate real-world error scenarios, in contrast to the staged mistakes of the current mistake-aware video datasets, and outperforms freeform LLM-based generation in coherence and perceived realism. Third, the thesis presents a lightweight, language-grounded PCMD framework, \texttt{ChronoFix}. The method grounds video executions into step sequences, compares raw step descriptions, semantic role representations, and action--object abstractions, and verifies the resulting traces with a Hidden Markov Model. Across CaptainCook4D, EgoPER, EgoOops, and auxiliary Assembly101 experiments, the results show that semantic-role normalization improves robustness to noisy VLM grounding and that explicit sequence modeling supports interpretable cross-dataset mistake detection. This work advances the state of the art by (1) providing diagnostic evidence of VLM failures in temporal logic, (2) introducing a scalable pipeline for generating realistic mistakes, and (3) presenting an efficient, structure-first baseline for post-completion mistake detection.</dc:description>
<dc:date>2026-04-30</dc:date>
<dc:type>info:eu-repo/semantics/doctoralThesis</dc:type>
<dc:identifier>https://hdl.handle.net/11572/484810</dc:identifier>
<dc:language>eng</dc:language>
<dc:relation>firstpage:1</dc:relation>
<dc:relation>lastpage:184</dc:relation>
<dc:relation>numberofpages:184</dc:relation>
<dc:rights>info:eu-repo/semantics/openAccess</dc:rights>
<dc:publisher>Università degli studi di Trento</dc:publisher>
<dc:publisher>place:TRENTO</dc:publisher>
<dc:rights>license:Tutti i diritti riservati (All rights reserved)</dc:rights>
<dc:rights>license uri:iris.PRI01</dc:rights>
</oai_dc:dc></metadata></record></GetRecord></OAI-PMH>