Browse State-of-the-Art › cross-modal alignment › Papers, page 3
cross-modal alignment
Papers archive 2025-07-28
archive papers tagged: 342 · with a code link: 151 · where Syntology ran a sample: 47 (41 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (47 of 342 tagged: 41 with a run with no instrument failure, 6 where every run was a failure of Syntology's instrument)
Page 3 of 4: papers 201 to 300 of 342, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
LangBridge: Interpreting Image as a Combination of Language Embeddings25 Mar 2025 0 repositories listed
-
Shushing! Let's Imagine an Authentic Speech from the Silent Video19 Mar 2025 0 repositories listed
-
Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation14 Mar 2025 0 repositories listed
-
Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition13 Mar 2025 0 repositories listed
-
4D-ACFNet: A 4D Attention Mechanism-Based Prognostic Framework for Colorectal Cancer Liver Metastasis Integrating Multimodal Spatiotemporal Features12 Mar 2025 0 repositories listed
-
Hierarchical Cross-Modal Alignment for Open-Vocabulary 3D Object Detection10 Mar 2025 0 repositories listed
-
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?10 Mar 2025 0 repositories listed
-
Enhancing Vision-Language Compositional Understanding with Multimodal Synthetic Data3 Mar 2025 0 repositories listed
-
Language Model Mapping in Multimodal Music Learning: A Grand Challenge Proposal1 Mar 2025 0 repositories listed
-
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting25 Feb 2025 0 repositories listed
-
DUNIA: Pixel-Sized Embeddings via Cross-Modal Alignment for Earth Observation Applications24 Feb 2025 0 repositories listed
-
CardiacMamba: A Multimodal RGB-RF Fusion Framework with State Space Models for Remote Physiological Measurement19 Feb 2025 0 repositories listed
-
A Survey of Automatic Prompt Engineering: An Optimization Perspective17 Feb 2025 0 repositories listed
-
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model17 Feb 2025 0 repositories listed
-
MDE: Modality Discrimination Enhancement for Multi-modal Recommendation8 Feb 2025 0 repositories listed
-
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion7 Feb 2025 0 repositories listed
-
Cross-modal Context Fusion and Adaptive Graph Convolutional Network for Multimodal Conversational Emotion Recognition25 Jan 2025 0 repositories listed
-
Integrate Temporal Graph Learning into LLM-based Temporal Knowledge Graph Model21 Jan 2025 0 repositories listed
-
CGP-Tuning: Structure-Aware Soft Prompt Tuning for Code Vulnerability Detection8 Jan 2025 0 repositories listed
-
Audio-Visual Semantic Graph Network for Audio-Visual Event Localization1 Jan 2025 0 repositories listed
-
Chat-based Person Retrieval via Dialogue-Refined Cross-Modal Alignment1 Jan 2025 0 repositories listed
-
Generalized Zero-Shot Classification via Semantics-Free Inter-Class Feature Generation1 Jan 2025 0 repositories listed
-
ChartAdapter: Large Vision-Language Model for Chart Summarization30 Dec 2024 0 repositories listed
-
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment30 Dec 2024 0 repositories listed
-
Bag of Tricks for Multimodal AutoML with Image, Text, and Tabular Data19 Dec 2024 0 repositories listed
-
RAC3: Retrieval-Augmented Corner Case Comprehension for Autonomous Driving with Vision-Language Models15 Dec 2024 0 repositories listed
-
Wearable Accelerometer Foundation Models for Health via Knowledge Distillation15 Dec 2024 0 repositories listed
-
Dynamic Cross-Modal Alignment for Robust Semantic Location Prediction13 Dec 2024 0 repositories listed
-
Enhancing Modality Representation and Alignment for Multimodal Cold-start Active Learning12 Dec 2024 0 repositories listed
-
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning10 Dec 2024 0 repositories listed
-
Towards Brain Passage Retrieval -- An Investigation of EEG Query Representations9 Dec 2024 0 repositories listed
-
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance5 Dec 2024 0 repositories listed
-
AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment1 Dec 2024 0 repositories listed
-
Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion27 Nov 2024 0 repositories listed
-
Tiny-Align: Bridging Automatic Speech Recognition and Large Language Model on the Edge21 Nov 2024 0 repositories listed
-
CTPD: Cross-Modal Temporal Pattern Discovery for Enhanced Multimodal Electronic Health Records Analysis1 Nov 2024 0 repositories listed
-
Multi-path Exploration and Feedback Adjustment for Text-to-Image Person Retrieval26 Oct 2024 0 repositories listed
-
Let Me Finish My Sentence: Video Temporal Grounding with Holistic Text Understanding17 Oct 2024 0 repositories listed
-
Modeling the Human Visual System: Comparative Insights from Response-Optimized and Task-Optimized Vision Models, Language Models, and different Readout Mechanisms17 Oct 2024 0 repositories listed
-
OMCAT: Omni Context Aware Transformer15 Oct 2024 0 repositories listed
-
EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment8 Oct 2024 0 repositories listed
-
Intriguing Properties of Large Language and Vision Models7 Oct 2024 0 repositories listed
-
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation5 Oct 2024 0 repositories listed
-
Fully Aligned Network for Referring Image Segmentation29 Sep 2024 0 repositories listed
-
Exploring Information-Theoretic Metrics Associated with Neural Collapse in Supervised Training25 Sep 2024 0 repositories listed
-
TS-HTFA: Advancing Time Series Forecasting via Hierarchical Text-Free Alignment with Large Language Models23 Sep 2024 0 repositories listed
-
Learning to Localize Actions in Instructional Videos with LLM-Based Multi-Pathway Text-Video Alignment22 Sep 2024 0 repositories listed
-
OneEncoder: A Lightweight Framework for Progressive Alignment of Modalities17 Sep 2024 0 repositories listed
-
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training15 Sep 2024 0 repositories listed
-
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization12 Sep 2024 0 repositories listed
-
GALLa: Graph Aligned Large Language Models for Improved Source Code Understanding6 Sep 2024 0 repositories listed
-
Temporal Order Preserved Optimal Transport-based Cross-modal Knowledge Transfer Learning for ASR3 Sep 2024 0 repositories listed
-
Coarse-to-fine Alignment Makes Better Speech-image Retrieval15 Aug 2024 0 repositories listed
-
Cross-Modal Denoising: A Novel Training Paradigm for Enhancing Speech-Image Retrieval15 Aug 2024 0 repositories listed
-
Cross-aware Early Fusion with Stage-divided Vision and Language Transformer Encoders for Referring Image Segmentation14 Aug 2024 0 repositories listed
-
Disentangled Noisy Correspondence Learning10 Aug 2024 0 repositories listed
-
Unifying Visual and Semantic Feature Spaces with Diffusion Models for Enhanced Cross-Modal Alignment26 Jul 2024 0 repositories listed
-
Multimodal Machine Learning in Mental Health: A Survey of Data, Algorithms, and Challenges23 Jul 2024 0 repositories listed
-
Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework12 Jul 2024 0 repositories listed
-
EA-VTR: Event-Aware Video-Text Retrieval10 Jul 2024 0 repositories listed
-
Cross-Modal Attention Alignment Network with Auxiliary Text Description for zero-shot sketch-based image retrieval1 Jul 2024 0 repositories listed
-
MLLM as Video Narrator: Mitigating Modality Imbalance in Video Moment Retrieval25 Jun 2024 0 repositories listed
-
Hire: Hybrid-modal Interaction with Multiple Relational Enhancements for Image-Text Matching5 Jun 2024 0 repositories listed
-
Multimodal Reasoning with Multimodal Knowledge Graph4 Jun 2024 0 repositories listed
-
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All25 May 2024 0 repositories listed
-
23 May 2024 0 repositories listed
-
Context-Enhanced Video Moment Retrieval with Large Language Models21 May 2024 0 repositories listed
-
Distributionally Robust Alignment for Medical Federated Vision-Language Pre-training Under Data Heterogeneity5 Apr 2024 0 repositories listed
-
CIRP: Cross-Item Relational Pre-training for Multimodal Product Bundling2 Apr 2024 0 repositories listed
-
Multi-Grained Cross-modal Alignment for Learning Open-vocabulary Semantic Segmentation from Text Supervision6 Mar 2024 0 repositories listed
-
Multi-modal Attribute Prompting for Vision-Language Models1 Mar 2024 0 repositories listed
-
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training1 Mar 2024 0 repositories listed
-
Mind the Modality Gap: Towards a Remote Sensing Vision-Language Model via Cross-modal Alignment15 Feb 2024 0 repositories listed
-
Cross-Modal Prototype based Multimodal Federated Learning under Severely Missing Modality25 Jan 2024 0 repositories listed
-
Multi-level Cross-modal Alignment for Image Clustering22 Jan 2024 0 repositories listed
-
6 Jan 2024 0 repositories listed Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Multi-Prompts Learning with Cross-Modal Alignment for Attribute-based Person Re-Identification28 Dec 2023 0 repositories listed
-
Detection-based Intermediate Supervision for Visual Question Answering26 Dec 2023 0 repositories listed
-
Improving Cross-modal Alignment with Synthetic Pairs for Text-only Image Captioning14 Dec 2023 0 repositories listed
-
OpenSight: A Simple Open-Vocabulary Framework for LiDAR-Based Object Detection12 Dec 2023 0 repositories listed
-
PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features5 Dec 2023 0 repositories listed
-
DAP: Domain-aware Prompt Learning for Vision-and-Language Navigation29 Nov 2023 0 repositories listed
-
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval30 Oct 2023 0 repositories listed
-
Video Referring Expression Comprehension via Transformer with Content-conditioned Query25 Oct 2023 0 repositories listed
-
On the Language Encoder of Contrastive Cross-modal Models20 Oct 2023 0 repositories listed
-
Robust Graph Matching Using An Unbalanced Hierarchical Optimal Transport Framework18 Oct 2023 0 repositories listed
-
Separating Invisible Sounds Toward Universal Audiovisual Scene-Aware Sound Separation18 Oct 2023 0 repositories listed
-
Prototype-guided Cross-modal Completion and Alignment for Incomplete Text-based Person Re-identification29 Sep 2023 0 repositories listed
-
Cross-modal Alignment with Optimal Transport for CTC-based ASR24 Sep 2023 0 repositories listed
-
Sound Source Localization is All about Cross-Modal Alignment19 Sep 2023 0 repositories listed
-
Prompt-based Context- and Domain-aware Pretraining for Vision and Language Navigation7 Sep 2023 0 repositories listed
-
Towards High-Fidelity Text-Guided 3D Face Generation and Manipulation Using only Images31 Aug 2023 0 repositories listed
-
DiffCloth: Diffusion Based Garment Synthesis and Manipulation via Structural Cross-modal Semantic Alignment22 Aug 2023 0 repositories listed
-
WiCo: Win-win Cooperation of Bottom-up and Top-down Referring Image Segmentation19 Jun 2023 0 repositories listed
-
15 Jun 2023 0 repositories listed
-
Improving speech translation by fusing speech and text23 May 2023 0 repositories listed
-
Multi-task Paired Masking with Alignment Modeling for Medical Vision-Language Pre-training13 May 2023 0 repositories listed
-
AlignSTS: Speech-to-Singing Conversion via Cross-Modal Alignment8 May 2023 0 repositories listed
-
CoVLR: Coordinating Cross-Modal Consistency and Intra-Modal Structure for Vision-Language Retrieval15 Apr 2023 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.