Time matters! capturing variation in time in video using Fisher Kernels

Mironica, I.; Uijlings, Jasper Reinout Robertus; Rostamzadeh, Negar; Ionescu, B.; Sebe, Niculae

doi:10.1145/2502081.2502183

In video global features are often used for reasons of computational efficiency, where each global feature captures information of a single video frame. But frames in video change over time, so an important question is: how can we meaningfully aggregate frame-based features in order to preserve the variation in time? In this paper we propose to use the Fisher Kernel to capture variation in time in video. While in this approach the temporal order is lost, it captures both subtle variation in time such as the ones caused by a moving bicycle and drastic variations in time such as the changing of shots in a documentary. Our work should not be confused with a Bag of Local Vi- sual Features approach, where one captures the visual variation of local features in both time and space indiscriminately. Instead, each feature measures a complete frame hence we capture variation in time only. We show that our framework is highly general, reporting improvements using frame-based visual features, body...