| Make Your Training Flexible: Towards Deployment-Efficient Video Models |
1 |
2 |
18 Mar 2025 |
official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 0 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (6 pointer-only for licence) |
| Gramian Multimodal Representation Learning and Alignment |
2 |
2 |
16 Dec 2024 |
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 1 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence) |
| vid-TLDR: Training Free Token merging for Light-weight Video Transformer |
1 |
2 |
20 Mar 2024 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization |
1 |
1 |
5 Feb 2024 |
3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence) |
| Multi-granularity Correspondence Learning from Long-term Noisy Videos |
1 |
1 |
30 Jan 2024 |
2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 2 samples that ran constructed an object rather than computing a result |
| MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing |
1 |
1 |
29 Nov 2023 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 2 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (12 pointer-only for licence) |
| HowToCaption: Prompting LLMs to Transform Video Annotations at Scale |
1 |
3 |
7 Oct 2023 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence) |
| LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment |
6 |
2 |
3 Oct 2023 |
official (archive's flag): 5 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (3 pointer-only for licence) |
| Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval |
1 |
1 |
29 Sep 2023 |
official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 1 violated, 7 with no contract checked; 7 where Syntology's instrument failed) · 2 unverified (8 pointer-only for licence) |
| BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning |
1 |
1 |
27 Sep 2023 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation |
1 |
1 |
27 Sep 2023 |
official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| Accurate and Fast Compressed Video Captioning |
1 |
1 |
22 Sep 2023 |
official (archive's flag): 10 ran · 10 ran (of which 8 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified |
| Unified Coarse-to-Fine Alignment for Video-Text Retrieval |
1 |
2 |
18 Sep 2023 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (4 pointer-only for licence) |
| ModelScope Text-to-Video Technical Report |
5 |
1 |
12 Aug 2023 |
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence) |
| VideoComposer: Compositional Video Synthesis with Motion Controllability |
4 |
1 |
3 Jun 2023 |
16 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (1 pointer-only for licence) |
| VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset |
2 |
3 |
29 May 2023 |
official (archive's flag): 12 ran · 35 ran (of which 4 constructed an object rather than computing a result; 29 with no instrument failure: 2 honoured, 1 violated, 26 with no contract checked; 6 where Syntology's instrument failed) · 7 unverified (8 pointer-only for licence) |
| ImageBind: One Embedding Space To Bind Them All |
3 |
1 |
9 May 2023 |
official (archive's flag): 18 ran · 24 ran (of which 14 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 1 violated, 19 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (32 pointer-only for licence) |
| Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models |
4 |
2 |
18 Apr 2023 |
official: no sample here; runs from other or unrecorded repositories · 18 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 3 violated, 7 with no contract checked; 6 where Syntology's instrument failed) · 8 unverified (3 pointer-only for licence) |
| MaMMUT: A Simple Architecture for Joint Learning for MultiModal Tasks |
1 |
1 |
29 Mar 2023 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| Unmasked Teacher: Towards Training-Efficient Video Foundation Models |
1 |
2 |
28 Mar 2023 |
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence) |
| Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning |
4 |
1 |
25 Mar 2023 |
official (archive's flag): 1 ran · 12 ran (of which 7 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified |
| MELTR: Meta Loss Transformer for Learning to Fine-tune Video Foundation Models |
1 |
7 |
23 Mar 2023 |
official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result |
| DiffusionRet: Generative Text-Video Retrieval with Diffusion Model |
4 |
2 |
17 Mar 2023 |
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified |
| mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video |
4 |
3 |
1 Feb 2023 |
community repositories only · 17 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence) |
| Cap4Video: What Can Auxiliary Captions Do for Text-Video Retrieval? |
4 |
1 |
31 Dec 2022 |
official (archive's flag): 14 ran · 20 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 2 honoured, 1 violated, 9 with no contract checked; 8 where Syntology's instrument failed) · 5 unverified (11 pointer-only for licence) |
| InternVideo: General Video Foundation Models via Generative and Discriminative Learning |
2 |
2 |
6 Dec 2022 |
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified |
| X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks |
2 |
2 |
22 Nov 2022 |
official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (6 pointer-only for licence) |
| Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations |
4 |
3 |
21 Nov 2022 |
official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified |
| CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment |
1 |
1 |
14 Sep 2022 |
official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence) |
| TS2-Net: Token Shift and Selection Transformer for Text-Video Retrieval |
1 |
1 |
16 Jul 2022 |
official (archive's flag): 4 ran · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (4 pointer-only for licence) |
| X-CLIP: End-to-End Multi-grained Contrastive Learning for Video-Text Retrieval |
3 |
1 |
15 Jul 2022 |
official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence) |
| Revealing Single Frame Bias for Video-and-Language Learning |
2 |
3 |
7 Jun 2022 |
community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| GIT: A Generative Image-to-text Transformer for Vision and Language |
1 |
1 |
27 May 2022 |
official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 0 honoured, 0 violated, 14 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified |
| CoCa: Contrastive Captioners are Image-Text Foundation Models |
6 |
1 |
4 May 2022 |
10 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 2 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (2 pointer-only for licence) |
| X-Pool: Cross-Modal Language-Video Attention for Text-Video Retrieval |
1 |
1 |
28 Mar 2022 |
official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (2 pointer-only for licence) |
| Disentangled Representation Learning for Text-Video Retrieval |
2 |
1 |
14 Mar 2022 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 2 unverified (4 pointer-only for licence) |
| Bridging Video-text Retrieval with Multiple Choice Questions |
2 |
3 |
13 Jan 2022 |
official (archive's flag): 2 ran · 13 ran (of which 8 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 11 unverified (6 pointer-only for licence) |
| Cross Modal Retrieval with Querybank Normalisation |
1 |
1 |
23 Dec 2021 |
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified |
| NÜWA: Visual Synthesis Pre-training for Neural visUal World creAtion |
1 |
1 |
24 Nov 2021 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence) |
| VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text |
5 |
1 |
22 Apr 2021 |
community repositories only · 5 ran (of which 4 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence) |
| CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval |
5 |
4 |
18 Apr 2021 |
official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval |
5 |
3 |
1 Apr 2021 |
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence) |
| MDMMT: Multidomain Multimodal Transformer for Video Retrieval |
3 |
2 |
19 Mar 2021 |
official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (5 pointer-only for licence) |
| Multi-modal Transformer for Video Retrieval |
1 |
3 |
21 Jul 2020 |
7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified |
| UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation |
2 |
1 |
15 Feb 2020 |
official: harvested, nothing ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (1 pointer-only for licence) |
| End-to-End Learning of Visual Representations from Uncurated Instructional Videos |
4 |
1 |
13 Dec 2019 |
official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence) |
| Use What You Have: Video Retrieval Using Representations From Collaborative Experts |
3 |
2 |
31 Jul 2019 |
3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence) |
| HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips |
4 |
3 |
7 Jun 2019 |
1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence) |
| A Joint Sequence Fusion Model for Video Question Answering and Retrieval |
2 |
2 |
7 Aug 2018 |
1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified |