Browse State-of-the-Art › cross-modal alignment › Papers, page 4
cross-modal alignment
Papers archive 2025-07-28
archive papers tagged: 342 · with a code link: 151 · where Syntology ran a sample: 47 (41 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (47 of 342 tagged: 41 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument)
Page 4 of 4: papers 301 to 342 of 342, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
SoftCLIP: Softer Cross-modal Alignment Makes CLIP Stronger30 Mar 2023 0 repositories listed
-
TOT: Topology-Aware Optimal Transport For Multimodal Hate Detection27 Feb 2023 0 repositories listed
-
End-to-end Semantic Object Detection with Cross-Modal Alignment10 Feb 2023 0 repositories listed
-
Does Vision Accelerate Hierarchical Generalization in Neural Language Learners?1 Feb 2023 0 repositories listed
-
Improving Cross-modal Alignment for Text-Guided Image Inpainting26 Jan 2023 0 repositories listed
-
Linguistic Query-Guided Mask Generation for Referring Image Segmentation16 Jan 2023 0 repositories listed
-
30 Dec 2022 0 repositories listed
-
7 Dec 2022 0 repositories listed
-
How do Cross-View and Cross-Modal Alignment Affect Representations in Contrastive Learning?23 Nov 2022 0 repositories listed
-
SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training21 Nov 2022 0 repositories listed
-
Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision24 Oct 2022 0 repositories listed
-
Fine-grained Semantic Alignment Network for Weakly Supervised Temporal Language Grounding21 Oct 2022 0 repositories listed
-
Cross-modal Semantic Enhanced Interaction for Image-Sentence Retrieval17 Oct 2022 0 repositories listed
-
Video Referring Expression Comprehension via Transformer with Content-aware Query6 Oct 2022 0 repositories listed
-
JPG - Jointly Learn to Align: Automated Disease Prediction and Radiology Report Generation1 Oct 2022 0 repositories listed
-
TokenFlow: Rethinking Fine-grained Cross-modal Alignment in Vision-Language Retrieval28 Sep 2022 0 repositories listed
-
28 Sep 2022 0 repositories listed
-
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval23 Sep 2022 0 repositories listed
-
15 Sep 2022 0 repositories listed
-
See What You See: Self-supervised Cross-modal Retrieval of Visual Stimuli from Brain Activity7 Aug 2022 0 repositories listed
-
Masked Vision and Language Modeling for Multi-modal Representation Learning3 Aug 2022 0 repositories listed
-
Cross-Modal Alignment Learning of Vision-Language Conceptual Systems31 Jul 2022 0 repositories listed
-
VLMixer: Unpaired Vision-Language Pre-training via Cross-Modal CutMix17 Jun 2022 0 repositories listed
-
mSLAM: Massively multilingual joint pre-training for speech and text3 Feb 2022 0 repositories listed
-
KD-VLP: Improving End-to-End Vision-and-Language Pretraining with Object Knowledge Distillation16 Jan 2022 0 repositories listed
-
Learning Better Visual Representations for Weakly-Supervised Object Detection Using Natural Language Supervision29 Sep 2021 0 repositories listed
-
Learning Joint Embedding with Modality Alignments for Cross-Modal Retrieval of Recipes and Food Images9 Aug 2021 0 repositories listed
-
Structured Multi-modal Feature Embedding and Alignment for Image-Sentence Retrieval5 Aug 2021 0 repositories listed
-
Continual learning in cross-modal retrieval14 Apr 2021 0 repositories listed
-
Scene-Intuitive Agent for Remote Embodied Visual Grounding24 Mar 2021 0 repositories listed
-
ST-BERT: Cross-modal Language Model Pre-training For End-to-end Spoken Language Understanding23 Oct 2020 0 repositories listed
-
Reinforcement Learning for Weakly Supervised Temporal Grounding of Natural Language in Untrimmed Videos18 Sep 2020 0 repositories listed
-
Cross-Modal Alignment with Mixture Experts Neural Network for Intral-City Retail Recommendation17 Sep 2020 0 repositories listed
-
Learning Multi-Modal Nonlinear Embeddings: Performance Bounds and an Algorithm3 Jun 2020 0 repositories listed
-
Cross-Modal Cross-Domain Moment Alignment Network for Person Search1 Jun 2020 0 repositories listed
-
Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models15 May 2020 0 repositories listed
-
11 May 2020 0 repositories listed
-
MCQA: Multimodal Co-attention Based Network for Question Answering25 Apr 2020 0 repositories listed
-
Curriculum Audiovisual Learning26 Jan 2020 0 repositories listed
-
ACMM: Aligned Cross-Modal Memory for Few-Shot Image and Sentence Matching1 Oct 2019 0 repositories listed
-
Mix and match networks: cross-modal alignment for zero-pair image-to-image translation8 Mar 2019 0 repositories listed
-
Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces18 May 2018 0 repositories listed