Browse State-of-the-Art › Self-supervised Video Retrieval
Self-supervised Video Retrieval
9 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
9 shown of 9 papers with code (10 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Dec 2022 2 repositories listedA good data representation should contain relations between the instances, or semantic similarity and dissimilarity, that contrastive learning harms by considering all negatives as noise.
-
6 Aug 2020 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)With the proposed Inter-Intra Contrastive (IIC) framework, we can train spatio-temporal convolutional networks to learn video representations.
-
25 Jun 2022 1 repository listedOne of the key reasons for this is that sampling pairs of similar video clips, a required step for many self-supervised contrastive learning methods, is currently done conservatively to avoid false positives.
-
9 Nov 2021 1 repository listedWe present CrissCross, a self-supervised framework for learning audio-visual representations.
-
18 Jun 2021 1 repository listedInstance-level contrastive learning techniques, which rely on data augmentation and a contrastive loss function, have found great success in the domain of visual representation learning.
-
20 Jan 2021 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedHowever, prior work on contrastive learning for video data has not explored the effect of explicitly encouraging the features to be distinct across the temporal dimension.
-
29 Oct 2020 1 repository listedIt is convenient to treat PCL as a standard training strategy and apply it to many other works in self-supervised video feature learning.
-
1 Jun 2020 1 repository listedThe generative perception model acts as a feature decoder to focus on comprehending high temporal resolution and short-term representation by introducing a motion-attention mechanism.
-
2 Jan 2020 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)As a proxy task, it converts rich self-supervised representations into video clip operations (options), which enhances the flexibility and reduces the complexity of representation learning.
Syntology lines on 3 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections