Papers › AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures

AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures

30 May 2019ICLR 2020 1arXiv:1905.13209archive 2025-07-28

Michael S. Ryoo, AJ Piergiovanni, Mingxing Tan, Anelia Angelova

Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to include the time dimension, using modules such as 3D convolutions, or by using two-stream design to capture both appearance and motion in videos. We interpret a video CNN as a collection of multi-stream convolutional blocks connected to each other, and propose the approach of automatically finding neural architectures with better connectivity and spatio-temporal interactions for video understanding. This is done by evolving a population of overly-connected architectures guided by connection weight learning. Architectures combining representations that abstract different input types (i.e., RGB and optical flow) at multiple temporal resolutions are searched for, allowing different types or sources of information to interact with each other. Our method, referred to as AssembleNet, outperforms prior approaches on public video datasets, in some cases by a great margin. We obtain 58.6% mAP on Charades and 34.27% accuracy on Moments-in-Time.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionMultimodal Activity RecognitionOptical Flow EstimationVideo ClassificationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Charades AssembleNet MAP 58.6 #6 of 49 Archive leaderboard report
Action Classification Charades AssembleNet-101 MAP 58.6 #7 of 49 Archive leaderboard report
Action Classification MiT AssembleNet Top 1 Accuracy 34.27% #16 of 29 Archive leaderboard report
Action Classification MiT AssembleNet Top 5 Accuracy 62.71% #16 of 29 Archive leaderboard report
Multimodal Activity Recognition Moments in Time Dataset AssembleNet Top-1 (%) 34.27 #1 of 6 Archive leaderboard report
Multimodal Activity Recognition Moments in Time Dataset AssembleNet Top-5 (%) 62.71 #1 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections