| Gramian Multimodal Representation Learning and Alignment |
2 |
2 |
16 Dec 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 1 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence) |
| vid-TLDR: Training Free Token merging for Light-weight Video Transformer |
1 |
2 |
20 Mar 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding |
1 |
1 |
14 Mar 2024 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence) |
| Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy |
1 |
1 |
11 Feb 2024 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding |
1 |
1 |
29 Oct 2023 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (5 pointer-only for licence) |
| LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment |
6 |
2 |
3 Oct 2023 |
official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence) |
| BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning |
1 |
1 |
27 Sep 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
1 |
29 May 2023 |
official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence) |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models |
1 |
2 |
28 Mar 2023 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning |
4 |
1 |
25 Mar 2023 |
official (archive's flag): 1 ran · 12 ran (of which 7 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified |
| DiffusionRet: Generative Text-Video Retrieval with Diffusion Model |
4 |
2 |
17 Mar 2023 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified |
| TriDet: Temporal Action Detection with Relative Boundary Modeling |
1 |
1 |
13 Mar 2023 |
official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (12 pointer-only for licence) |
| InternVideo: General Video Foundation Models via Generative and Discriminative Learning |
2 |
3 |
6 Dec 2022 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified |
| Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations |
4 |
2 |
21 Nov 2022 |
official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment |
1 |
1 |
14 Sep 2022 |
official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence) |
| X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval |
3 |
1 |
15 Jul 2022 |
official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| Proposal-Free Temporal Action Detection via Global Segmentation Mask Learning |
2 |
1 |
14 Jul 2022 |
official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (3 pointer-only for licence) |
| Revealing Single Frame Bias for Video-and-Language Learning |
2 |
3 |
7 Jun 2022 |
community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| An Empirical Study of End-to-End Temporal Action Detection |
1 |
1 |
6 Apr 2022 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (11 pointer-only for licence) |
| Audio-visual Generalised Zero-shot Learning with Cross-modal Attention and Language |
1 |
2 |
7 Mar 2022 |
official (archive's flag): 3 ran · 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 3 samples that ran constructed an object rather than computing a result |
| End-to-end Temporal Action Detection with Transformer |
1 |
1 |
18 Jun 2021 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified |
| DSANet: Dynamic Segment Aggregation Network for Video-Level Representation Learning |
1 |
1 |
25 May 2021 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified |
| Weakly Supervised Action Selection Learning in Video |
1 |
1 |
6 May 2021 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval |
5 |
1 |
18 Apr 2021 |
official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| Video Self-Stitching Graph Network for Temporal Action Localization |
1 |
1 |
30 Nov 2020 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| BSN++: Complementary Boundary Regressor with Scale-Balanced Relation Modeling for Temporal Action Proposal Generation |
1 |
2 |
15 Sep 2020 |
9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| Multi-modal Transformer for Video Retrieval |
1 |
2 |
21 Jul 2020 |
7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| Rethinking Zero-shot Video Classification: End-to-end Training for Realistic Applications |
1 |
1 |
3 Mar 2020 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified |
| G-TAD: Sub-Graph Localization for Temporal Action Detection |
7 |
1 |
26 Nov 2019 |
official (archive's flag): 11 ran · 15 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (2 pointer-only for licence) |
| Use What You Have: Video Retrieval Using Representations From Collaborative Experts |
3 |
1 |
31 Jul 2019 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| BMN: Boundary-Matching Network for Temporal Action Proposal Generation |
15 |
2 |
23 Jul 2019 |
community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| W-TALC: Weakly-supervised Temporal Activity Localization and Classification |
1 |
2 |
27 Jul 2018 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| CTAP: Complementary Temporal Action Proposal Generation |
1 |
1 |
12 Jul 2018 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified |
| BSN: Boundary Sensitive Network for Temporal Action Proposal Generation |
17 |
2 |
8 Jun 2018 |
community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (4 pointer-only for licence) |
| YouTube-8M: A Large-Scale Video Classification Benchmark |
7 |
1 |
27 Sep 2016 |
7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified |