{"url":"/method/sscs","slug":"sscs","name":"Sscs","full_name":"Support-set Based Cross-Supervision","full_name_withheld":false,"description_markdown":"**Sscs**, or **Support-set Based Cross-Supervision**, is a module for video grounding which consists of two main components: a discriminative contrastive objective and a generative caption objective. The contrastive objective aims to learn effective representations by contrastive learning, while the caption objective can train a powerful video encoder supervised by texts. Due to the co-existence of some visual entities in both ground-truth and background intervals, i.e., mutual exclusion, naively contrastive learning is unsuitable to video grounding. This problem is addressed by boosting the cross-supervision with the support-set concept, which collects visual information from the whole video and eliminates the mutual exclusion of entities.\r\n\r\nSpecifically, in the Figure to the right, two video-text pairs { $V\\_{i}, L\\_{i}$}, {$V\\_{j} , L\\_{j}$ } in the batch are presented for clarity. After feeding them into a video and text encoder, the clip-level and sentence-level embedding ( {$X\\_{i}, Y\\_{i}$} and {$X\\_{j} , Y\\_{j}$} ) in a shared space are acquired. Base on the support-set module, the weighted average of $X\\_{i}$ and $X\\_{j}$ is computed to obtain $\\bar{X}\\_{i}$, $\\bar{X}\\_{j}$ respectively. Finally, the contrastive and caption objectives are combined to pull close the representations of the clips and text from the same samples and push away those from other pairs","description_state":"present","introduced_year":null,"introduced_by":{"title":"Support-Set Based Cross-Supervision for Video Grounding","paper":"/paper/support-set-based-cross-supervision-for-video","first_author":"Xinpeng Ding","n_authors":8,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/support-set-based-cross-supervision-for-video"},"source":{"url":"https://arxiv.org/abs/2108.10576v1","title":"Support-Set Based Cross-Supervision for Video Grounding","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Video Model Blocks","url":"/methods/category/video-model-blocks","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/support-set-based-cross-supervision-for-video","title":"Support-Set Based Cross-Supervision for Video Grounding","date":"2021-08-24","arxiv_id":"2108.10576","n_code_links":0,"syntology":null}],"papers_shown":1,"tasks":[{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/video-grounding","name":"Video Grounding","papers":1}],"tasks_shown":2,"n_tasks":2,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sscs"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}