Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.
VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration / Yuan, H., Liu, Z., Zhou, J., Qian, H., Shu, Y., Sebe, N., Wen, J., Dou, Z.. - (2026), pp. 6350-6361. (ACM SIGKDD Conference on Knowledge Discovery and Data Mining Korea August 2026) [10.1145/3770855.3817705].
VideoExplorer: Advancing Long-Horizon Video Understanding via Hierarchical Orchestration
Shu, Yan;Sebe, Nicu;
2026-01-01
Abstract
Current agentic frameworks for Long-Video Understanding (LVU) remain limited by two critical problems: ineffective control, where traditional monolithic agents struggle with high-branching, multi-granularity decision processes; and inefficient supervision, where sparse, outcome-based feedback fails to guide long-horizon reasoning. To resolve these challenges, we propose VideoExplorer, a novel agentic system designed to advance long-video reasoning on top of structured control and trajectory-level optimization. First, VideoExplorer innovates a hierarchically orchestrated framework: it employs a planning agent to focus on creating high-level reasoning strategies and specialized sub-agents (including a temporal grounder and a visual perceiver) to accomplish fine-grained reasoning executions, thereby substantially reducing the complexity of reasoning process. Second, VideoExplorer introduces a novel optimization approach, Trajectory level Direct Preference Optimization (TDPO), to mitigate inefficient supervision. Unlike standard methods that optimize single turn responses, TDPO aligns the entire planning trajectory, including evidence routing and termination decisions, with end task success, which effectively mitigates premature commitment and compounding errors. To better support the conduct of TDPO, we further create fine-grained supervision data via a teacher-guided, difficulty-adaptive sampling process. Extensive experiments on MLVU, LVBench, and MH-NIAH demonstrate that VideoExplorer consistently outperforms monolithic baselines in both accuracy and efficiency, validating the effectiveness of structured control and trajectory-level optimization in long-video reasoning. Our code is available in this repository.| File | Dimensione | Formato | |
|---|---|---|---|
|
3770855.3817705-compressed.pdf
accesso aperto
Tipologia:
Versione editoriale (Publisher’s layout)
Licenza:
Creative commons
Dimensione
1.18 MB
Formato
Adobe PDF
|
1.18 MB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione



