Browse State-of-the-Art › Multi-modal Recommendation
Multi-modal Recommendation
18 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Amazon Baby (10 rows) | MIG-GT | Modality-Independent Graph Neural Networks with Global... | code | — | Compare |
| Amazon Clothing (10 rows) | MIG-GT | Modality-Independent Graph Neural Networks with Global... | code | — | Compare |
| Amazon Sports (10 rows) | MIG-GT | Modality-Independent Graph Neural Networks with Global... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
18 shown of 18 papers with code (30 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Feb 2020 18 repositories listed Syntology ran 2 of 7 samples · 5 unverified · 2 pointer-only (licence)We propose a new model named LightGCN, including only the most essential component in GCN -- neighborhood aggregation -- for collaborative filtering.
-
6 Oct 2015 7 repositories listed Syntology ran 3 of 28 samples · 25 unverified · 3 pointer-only (licence)In this paper we propose a scalable factorization model to incorporate visual signals into predictors of people's opinions, which we apply to a selection of large, real-world datasets.
-
21 Feb 2023 2 repositories listedThe online emergence of multi-modal sharing platforms (eg, TikTok, Youtube) is powering personalized recommender systems to incorporate various modalities (eg, visual, textual and acoustic) into the latent user…
-
13 Nov 2022 2 repositories listedBased on this finding, we propose a simple yet effective model, dubbed as FREEDOM, that FREEzes the item-item graph and DenOises the user-item interaction graph simultaneously for Multimodal recommendation.
-
13 Jul 2022 2 repositories listedBesides the user-item interaction graph, existing state-of-the-art methods usually use auxiliary graphs (e.
-
5 Apr 2022 2 repositories listedBased on a Markov process that trades off two types of distances, we present Markov Graph Diffusion Collaborative Filtering (MGDCF) to generalize some state-of-the-art GNN-based CF models.
-
25 Jun 2025 1 repository listedThis paper addresses the challenge of developing multimodal recommender systems for the movie domain, where limited metadata (e.
-
19 Apr 2025 1 repository listedHowever, MMRecs struggle with noisy data caused by misalignment among modal content and the gap between modal semantics and recommendation semantics.
-
18 Dec 2024 1 repository listedOur results indicate that the optimal K for certain modalities on specific datasets can be as low as 1 or 2, which may restrict the GNNs' capacity to capture global information.
-
19 Aug 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Recent advances in Large Language Models (LLMs) have demonstrated significant potential in the field of Recommendation Systems (RSs).
-
7 Jul 2024 1 repository listedMulti-modal recommendation greatly enhances the performance of recommender systems by modeling the auxiliary information from multi-modality contents.
-
16 Feb 2024 1 repository listedIn the feature extract phase, for image features, we are the first to combine image painting style features with semantic features to construct a dual-output image encoder for enhancing representation.
-
8 Aug 2023 1 repository listedTo be specific, we first introduce an ID-aware Multi-modal Transformer module in the item representation learning stage to facilitate information interaction among different features.
-
30 Jun 2022 1 repository listedTo capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation.
-
24 Dec 2021 1 repository listedSpecifically, we first introduce a single-modal representation learning module, which performs graph operations on the user-microvideo graph in each modality to capture single-modal user preferences on different…
-
3 Nov 2021 1 repository listedReorganizing implicit feedback of users as a user-item interaction graph facilitates the applications of graph convolutional networks (GCNs) in recommendation tasks.
-
19 Apr 2021 1 repository listed Syntology ran 3 of 3 samples · 0 unverifiedTo be specific, in the proposed LATTICE model, we devise a novel modality-aware structure learning layer, which learns item-item structures for each modality and aggregates multiple modalities to obtain latent item…
-
19 Oct 2019 1 repository listedExisting works on multimedia recommendation largely exploit multi-modal contents to enrich item representations, while less effort is made to leverage information interchange between users and items to enhance user…
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections