Papers › AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
AssembleNet: Searching for Multi-Stream Neural Connectivity in Video Architectures
Michael S. Ryoo, AJ Piergiovanni, Mingxing Tan, Anelia Angelova
Learning to represent videos is a very challenging task both algorithmically and computationally. Standard video CNN architectures have been designed by directly extending architectures devised for image understanding to include the time dimension, using modules such as 3D convolutions, or by using two-stream design to capture both appearance and motion in videos. We interpret a video CNN as a collection of multi-stream convolutional blocks connected to each other, and propose the approach of automatically finding neural architectures with better connectivity and spatio-temporal interactions for video understanding. This is done by evolving a population of overly-connected architectures guided by connection weight learning. Architectures combining representations that abstract different input types (i.e., RGB and optical flow) at multiple temporal resolutions are searched for, allowing different types or sources of information to interact with each other. Our method, referred to as AssembleNet, outperforms prior approaches on public video datasets, in some cases by a great margin. We obtain 58.6% mAP on Charades and 34.27% accuracy on Moments-in-Time.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Classification | Charades | AssembleNet | MAP | 58.6 | #6 of 49 | Archive leaderboard | report |
| Action Classification | Charades | AssembleNet-101 | MAP | 58.6 | #7 of 49 | Archive leaderboard | report |
| Action Classification | MiT | AssembleNet | Top 1 Accuracy | 34.27% | #16 of 29 | Archive leaderboard | report |
| Action Classification | MiT | AssembleNet | Top 5 Accuracy | 62.71% | #16 of 29 | Archive leaderboard | report |
| Multimodal Activity Recognition | Moments in Time Dataset | AssembleNet | Top-1 (%) | 34.27 | #1 of 6 | Archive leaderboard | report |
| Multimodal Activity Recognition | Moments in Time Dataset | AssembleNet | Top-5 (%) | 62.71 | #1 of 6 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections