| When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions added by Syntology |
2025-10 (from id) |
Zhuo-Cao/QV-M2/FlashMMR/position_encoding.py 1affbf903fa3204b |
unverified |
no licence file found · pointer only |
| arXiv:2507.19993 |
2025-07 (from id) |
Howardkhh/FROSS/EGTR/model/deformable_detr.py 72ceaeb9390c2de9 |
unverified |
Apache-2.0 (permissive) |
| Disambiguating Reference in Visually Grounded Dialogues through Joint Modeling of Textual and Multimodal Semantic Structures |
16 May 2025 |
ashkamath/mdetr/models/position_encoding.py cea3f92defccfb9c |
unverified |
Apache-2.0 (permissive) |
| VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary |
12 Mar 2025 |
showlab/VLog/VLog/model/models.py 020b7f70a236fb97 |
ran · our draft was wrong
|
no licence file found · pointer only |
| LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models |
31 Jan 2025 |
iSEE-Laboratory/LLMDet/hf_model/modeling_grounding_dino.py d23bffc29f903730 |
unverified |
Apache-2.0 recorded; this copy not marked cleared · pointer only |
| The Devil is in the Spurious Correlation: Boosting Moment Retrieval via Temporal Dynamic Learning |
13 Jan 2025 |
xyangzhou/TD-DETR/td_detr/position_encoding.py d6ca656c20df137a |
ran
|
MIT (permissive) |
| Interpretable Enzyme Function Prediction via Residue-Level Detection |
10 Jan 2025 |
yangzhao1230/protdetr/models/position_encoding.py ea84c6853134f470 |
unverified |
MIT (permissive) |
| FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding |
18 Dec 2024 |
zhuo-cao/flashvtg/FlashVTG/position_encoding.py 1affbf903fa3204b |
unverified |
no licence file found · pointer only |
| Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval |
21 Jul 2024 |
fletcherjiang/llmepet/llm_epet/position_encoding.py 79c51f8fc390ab83 |
ran
|
BSD-3-Clause (permissive) |
| SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding |
6 Jul 2024 |
zxccade/SHINE/shine/position_encoding.py d6ca656c20df137a |
ran
|
licence not identified · pointer only |
| SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding |
3 Jul 2024 |
WeitaiKang/SegVG/models/SegVG.py 153726d9d215fcbd |
ran · our draft was wrong
|
no licence file found · pointer only |
| Task-Driven Exploration: Decoupling and Inter-Task Feedback for Joint Moment Retrieval and Highlight Detection |
14 Apr 2024 |
EdenGabriel/TaskWeave/taskweave/position_encoding.py d6ca656c20df137a |
ran
|
no licence file found · pointer only |
| EGTR: Extracting Graph from Transformer for Scene Graph Generation |
2 Apr 2024 |
naver-ai/egtr/model/deformable_detr.py 72ceaeb9390c2de9 |
unverified |
Apache-2.0 (permissive) |
| A Simple Baseline for Efficient Hand Mesh Reconstruction |
4 Mar 2024 |
patiencefromzhou/simplehand/models/position_embedding.py e75bcdbf04de72fc |
unverified |
MIT (permissive) |
| TR-DETR: Task-Reciprocal Transformer for Joint Moment Retrieval and Highlight Detection |
4 Jan 2024 |
mingyao1120/tr-detr/tr_detr/position_encoding.py d6ca656c20df137a |
ran
|
no licence file found · pointer only |
| Context-Guided Spatio-Temporal Video Grounding |
3 Jan 2024 |
HengLan/CGSTVG/models/pipeline.py e936bfbe7328d661 |
ran · our draft was wrong
|
no licence file found · pointer only |
| BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos |
30 Nov 2023 |
Pilhyeon/BAM-DETR/bam_detr/position_encoding.py e8a160c4a8f899b5 |
ran
|
licence not identified · pointer only |
| Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection |
28 Nov 2023 |
easonxiao-888/uvcom/uvcom/position_encoding.py d6ca656c20df137a |
ran
|
MIT (permissive) |
| Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding |
15 Nov 2023 |
wjun0830/cgdetr/cg_detr/position_encoding.py 1affbf903fa3204b |
unverified |
no licence file found · pointer only |
| Shatter and Gather: Learning Referring Image Segmentation with Text Supervision |
29 Aug 2023 |
kdwonn/SaG/model/cross_modal_attention.py 1508f01cf4f224d9 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Point-Query Quadtree for Crowd Counting, Localization, and More |
26 Aug 2023 |
cxliu0/PET/models/pet.py f57a0aab388112d0 |
unverified |
MIT (permissive) |
| Knowing Where to Focus: Event-aware Transformer for Video Grounding |
14 Aug 2023 |
jinhyunj/eatr/models/position_encoding.py b4196f787bec2e3e |
ran
|
MIT (permissive) |
| Towards Generalist Foundation Model for Radiology by Leveraging Web-scale 2D&3D Medical Data |
4 Aug 2023 |
chaoyi-wu/radfm/Quick_demo/Model/RadFM/position_encoding.py c8f7b3c548c9d56a |
ran
|
MIT (permissive) |
| UniVTG: Towards Unified Video-Language Temporal Grounding |
31 Jul 2023 |
showlab/univtg/model/position_encoding.py 2912d21d62514265 |
ran
|
MIT (permissive) |
| MomentDiff: Generative Video Moment Retrieval from Random to Real |
6 Jul 2023 |
imccretrieval/momentdiff/momentdiff/position_encoding.py d6ca656c20df137a |
ran
|
no licence file found · pointer only |
| EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction |
30 Apr 2023 |
ercanburak/EVREAL/model/eitr/position_encoding.py bfc2d3b5c69ffb3b |
unverified |
MIT (permissive) |
| Learning Bottleneck Concepts in Image Classification |
20 Apr 2023 |
wbw520/botcl/model/reconstruct/model_main.py 59886ab4f3558c49 |
ran · our draft was wrong
|
no licence file found · pointer only |
| Sampling is Matter: Point-guided 3D Human Mesh Reconstruction |
19 Apr 2023 |
DCVL-3D/PointHMR_release/src/modeling/model/network.py c62bd3602bcf3ba2 |
ran · our draft was wrong
|
MIT (permissive) |
| Extending Phrase Grounding with Pronouns in Visual Dialogues |
23 Oct 2022 |
izhx/Phrase-Grounding-with-Pronoun/code/src/mdetr/position_encoding.py 1b6e04ba3ceda006 |
unverified |
Apache-2.0 (permissive) |
| Embracing Consistency: A One-Stage Approach for Spatio-Temporal Video Grounding |
27 Sep 2022 |
jy0205/stcat/models/pipeline.py ebf7ea1f0aba8315 |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection |
6 Sep 2022 |
Hon-Wong/PTSEFormer/src/models/model_builder.py f2cdbc05e138fb89 |
ran · our draft was wrong
|
MIT (permissive) |
| Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers |
27 Jul 2022 |
postech-ami/fastmetro/src/modeling/model/modeling_fastmetro.py 4407bd06305a2f6c |
ran · our draft was wrong
|
MIT (permissive) |
| SiRi: A Simple Selective Retraining Mechanism for Transformer-based Visual Grounding |
27 Jul 2022 |
qumengxue/siri-vg/models/position_encoding.py cea3f92defccfb9c |
unverified |
Apache-2.0 (permissive) |
| UniTAB: Unifying Text and Box Outputs for Grounded Vision-Language Modeling |
23 Nov 2021 |
microsoft/UniTAB/models/position_encoding.py cea3f92defccfb9c |
unverified |
MIT (permissive) |
| Transformer-based Dual Relation Graph for Multi-label Image Recognition |
10 Oct 2021 |
iCVTEAM/TDRG/models/TDRG.py 626421743e751cfd |
ran · our draft was wrong
|
Apache-2.0 (permissive) |
| Transformer for Single Image Super-Resolution |
25 Aug 2021 |
luissen/esrt/util/position.py 61eaa43ff60ccf54 |
unverified |
MIT (permissive) |
| MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding |
26 Apr 2021 |
b-faye/lightmdetr/models/position_encoding.py cea3f92defccfb9c |
unverified |
Apache-2.0 (permissive) |
| Video-aided Unsupervised Grammar Induction |
9 Apr 2021 |
Sy-Zhang/MMC-PCFG/lib/model/vpcfg/position_encoding.py ea9874d5c6513212 |
unverified |
MIT (permissive) |
| Learning Multi-Scene Absolute Pose Regression with Transformers |
21 Mar 2021 |
yolish/c2f-ms-transformer/models/transposenet/C2FEMSTransPoseNet.py 9cbd4a5c6cc2d670 |
ran · our draft was wrong
|
no licence file found · pointer only |
| arXiv:aaai_29295 |
|
Jack24658735/FedLGT/models/utils.py a809b63d1fb6ac5c |
unverified |
Apache-2.0 (permissive) |
| arXiv:aaai_27766 |
|
7tl7qns7ch/IPOT/models/ipot/position_encoding.py 92655046480cb7fb |
unverified |
MIT (permissive) |
| arXiv:Xiao_Bridging_the_Gap_A_Unified_Video_Comprehension_Framework_for_Moment_CVPR_2024_paper |
|
EasonXiao-888/UVCOM/uvcom/position_encoding.py d6ca656c20df137a |
ran
|
MIT (permissive) |