Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.

VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration / Yuan, H., Liu, Z., Zhou, J., Qian, H., Shu, Y., Sebe, N., Wen, J., Dou, Z.. - (2026), pp. 6350-6361. (ACM SIGKDD Conference on Knowledge Discovery and Data Mining Korea August 2026) [10.1145/3770855.3817705].

VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration

Shu, Yan;Sebe, Nicu;
2026-01-01

Abstract

Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.
2026
KDD '26: Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining
New York
Association for Computing Machinery, Inc
Yuan, Huaying; Liu, Zheng; Zhou, Junjie; Qian, Hongjin; Shu, Yan; Sebe, Nicu; Wen, Ji-Rong; Dou, Zhicheng
VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration / Yuan, H., Liu, Z., Zhou, J., Qian, H., Shu, Y., Sebe, N., Wen, J., Dou, Z.. - (2026), pp. 6350-6361. (ACM SIGKDD Conference on Knowledge Discovery and Data Mining Korea August 2026) [10.1145/3770855.3817705].
File in questo prodotto:
File Dimensione Formato  
3770855.3817705-compressed.pdf

accesso aperto

Tipologia: Versione editoriale (Publisher’s layout)
Licenza: Creative commons
Dimensione 1.18 MB
Formato Adobe PDF
1.18 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11572/498343
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact