Papers › DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition

DMC-Net: Generating Discriminative Motion Cues for Fast Compressed Video Action Recognition

11 Jan 2019CVPR 2019 6arXiv:1901.03460archive 2025-07-28

Zheng Shou, Xudong Lin, Yannis Kalantidis, Laura Sevilla-Lara, Marcus Rohrbach, Shih-Fu Chang, Zhicheng Yan

Motion has shown to be useful for video understanding, where motion is typically represented by optical flow. However, computing flow from video frames is very time-consuming. Recent works directly leverage the motion vectors and residuals readily available in the compressed video to represent motion at no cost. While this avoids flow computation, it also hurts accuracy since the motion vector is noisy and has substantially reduced resolution, which makes it a less discriminative motion representation. To remedy these issues, we propose a lightweight generator network, which reduces noises in motion vectors and captures fine motion details, achieving a more Discriminative Motion Cue (DMC) representation. Since optical flow is a more accurate motion representation, we train the DMC generator to approximate flow using a reconstruction loss and a generative adversarial loss, jointly with the downstream action classification task. Extensive evaluations on three action recognition benchmarks (HMDB-51, UCF-101, and a subset of Kinetics) confirm the effectiveness of our method. Our full system, consisting of the generator and the classifier, is coined as DMC-Net which obtains high accuracy close to that of using flow and runs two orders of magnitude faster than using optical flow at inference time.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAction Recognition In VideosOptical Flow EstimationTemporal Action LocalizationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition HMDB-51 I3D RGB + DMC-Net (I3D) Average accuracy of 3 splits 77.8 #30 of 77 Archive leaderboard report
Action Recognition HMDB-51 DMC-Net (I3D) Average accuracy of 3 splits 71.8 #50 of 77 Archive leaderboard report
Action Recognition HMDB-51 DMC-Net (ResNet-18) Average accuracy of 3 splits 62.8 #66 of 77 Archive leaderboard report
Action Recognition UCF-101 DMC-Net (ResNet-18) 3-fold Accuracy 90.9 #1 of 2 Archive leaderboard report
Action Recognition UCF101 I3D RGB + DMC-Net (I3D) 3-fold Accuracy 96.5 #37 of 91 Archive leaderboard report
Action Recognition UCF101 DMC-Net (I3D) 3-fold Accuracy 92.3 #64 of 91 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections