{"url":"/task/multimodal-recommendation","name":"Multimodal Recommendation","slug":"multimodal-recommendation","description_markdown":"The multimodal recommendation task involves developing systems that leverage and integrate multiple types of data—such as text, images, audio, and user interactions—to predict and suggest items that align with a user's preferences. Unlike traditional recommendation approaches that rely on a single data modality, multimodal recommendation harnesses the diverse information from various sources to create richer and more nuanced representations of both users and items. This integration enables the system to understand and capture complex relationships and attributes across different data types, thereby enhancing the accuracy and relevance of the recommendations. The primary goal is to provide personalized suggestions by effectively merging and processing heterogeneous data to better match users with items they are likely to engage with or find valuable.","categories":[{"name":"Graphs","url":"/area/graphs"},{"name":"Knowledge Base","url":"/area/knowledge-base"},{"name":"Miscellaneous","url":"/area/miscellaneous"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":59,"papers_with_code":33,"benchmarks":5,"benchmark_tables_in_archive":5,"benchmark_tables_shown":5,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":6,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/multimodal-recommendation-on-amazon-baby","slug":"multimodal-recommendation-on-amazon-baby","dataset":"Amazon Baby","dataset_url":"/dataset/amazon-baby","rows_in_archive":18,"metrics":["Recall","nDCG","Hit Ratio"],"first_row_in_archive_order":{"model":"FREEDOM (CLIP)","paper_title":"Ducho meets Elliot: Large-scale Benchmarks for Multimodal Recommendation","paper_url":"/paper/ducho-meets-elliot-large-scale-benchmarks-for-1","paper_date":"2024-09-24","arxiv_id":"2409.15857","code_links":[{"title":"sisinflab/Ducho-meets-Elliot","url":"https://github.com/sisinflab/Ducho-meets-Elliot"}],"syntology":null}},{"leaderboard":"/sota/multimodal-recommendation-on-amazon-beauty","slug":"multimodal-recommendation-on-amazon-beauty","dataset":"Amazon Beauty","dataset_url":"/dataset/amazon-beauty","rows_in_archive":18,"metrics":["Recall","nDCG","Hit Ratio"],"first_row_in_archive_order":{"model":"LATTICE (ALIGN)","paper_title":"Ducho meets Elliot: Large-scale Benchmarks for Multimodal Recommendation","paper_url":"/paper/ducho-meets-elliot-large-scale-benchmarks-for-1","paper_date":"2024-09-24","arxiv_id":"2409.15857","code_links":[{"title":"sisinflab/Ducho-meets-Elliot","url":"https://github.com/sisinflab/Ducho-meets-Elliot"}],"syntology":null}},{"leaderboard":"/sota/multimodal-recommendation-on-amazon-digital","slug":"multimodal-recommendation-on-amazon-digital","dataset":"Amazon Digital Music","dataset_url":"/dataset/amazon-digital-music","rows_in_archive":18,"metrics":["Recall","nDCG","Hit Ratio"],"first_row_in_archive_order":{"model":"LATTICE (AltCLIP)","paper_title":null,"paper_url":null,"paper_date":"","arxiv_id":null,"code_links":[],"syntology":null}},{"leaderboard":"/sota/multimodal-recommendation-on-amazon-office","slug":"multimodal-recommendation-on-amazon-office","dataset":"Amazon Office Products","dataset_url":"/dataset/amazon-office-products","rows_in_archive":18,"metrics":["Recall","nDCG","Hit Ratio"],"first_row_in_archive_order":{"model":"LATTICE (ResNet50+ Sentence Bert)","paper_title":"Ducho meets Elliot: Large-scale Benchmarks for Multimodal Recommendation","paper_url":"/paper/ducho-meets-elliot-large-scale-benchmarks-for-1","paper_date":"2024-09-24","arxiv_id":"2409.15857","code_links":[{"title":"sisinflab/Ducho-meets-Elliot","url":"https://github.com/sisinflab/Ducho-meets-Elliot"}],"syntology":null}},{"leaderboard":"/sota/multimodal-recommendation-on-amazon-toys","slug":"multimodal-recommendation-on-amazon-toys","dataset":"Amazon Toys & Games","dataset_url":"/dataset/amazon-toys-games","rows_in_archive":18,"metrics":["Recall","nDCG","Hit Ratio"],"first_row_in_archive_order":{"model":"FREEDOM (MMFashion + Sentence Bert)","paper_title":"Ducho meets Elliot: Large-scale Benchmarks for Multimodal Recommendation","paper_url":"/paper/ducho-meets-elliot-large-scale-benchmarks-for-1","paper_date":"2024-09-24","arxiv_id":"2409.15857","code_links":[{"title":"sisinflab/Ducho-meets-Elliot","url":"https://github.com/sisinflab/Ducho-meets-Elliot"}],"syntology":null}}],"datasets":[{"url":"/dataset/amazon-review","name":"Amazon Review","full_name":"","num_papers_in_archive":35},{"url":"/dataset/amazon-beauty","name":"Amazon Beauty","full_name":"Amazon Beauty 5-core","num_papers_in_archive":18},{"url":"/dataset/amazon-baby","name":"Amazon Baby","full_name":"Amazon Baby 5-core","num_papers_in_archive":13},{"url":"/dataset/amazon-toys-games","name":"Amazon Toys & Games","full_name":"Amazon Toys & Games 5-core","num_papers_in_archive":8},{"url":"/dataset/amazon-digital-music","name":"Amazon Digital Music","full_name":"Amazon Digital Music 5-core","num_papers_in_archive":3},{"url":"/dataset/amazon-office-products","name":"Amazon Office Products","full_name":"Amazon Office Products 5-core","num_papers_in_archive":3}],"subtasks":[],"parent_tasks":[{"url":"/task/recommendation-systems","name":"Recommendation Systems"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":33,"tagged_in_all":59,"items":[{"url":"/paper/a-comprehensive-survey-on-multimodal","title":"A Comprehensive Survey on Multimodal Recommender Systems: Taxonomy, Evaluation, and Future Directions","date":"2023-02-09","arxiv_id":"2302.04473","repositories_listed":2,"syntology":null},{"url":"/paper/a-tale-of-two-graphs-freezing-and-denoising","title":"A Tale of Two Graphs: Freezing and Denoising Graph Structures for Multimodal Recommendation","date":"2022-11-13","arxiv_id":"2211.06924","repositories_listed":2,"syntology":null},{"url":"/paper/quadratic-interest-network-for-multimodal","title":"Quadratic Interest Network for Multimodal Click-Through Rate Prediction","date":"2025-04-24","arxiv_id":"2504.17699","repositories_listed":1,"syntology":null},{"url":"/paper/cohesion-composite-graph-convolutional","title":"COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation","date":"2025-04-06","arxiv_id":"2504.04452","repositories_listed":1,"syntology":null},{"url":"/paper/collaborative-filtering-meets-spectrum-shift","title":"Collaborative Filtering Meets Spectrum Shift: Connecting User-Item Interaction with Graph-Structured Side Information","date":"2025-02-12","arxiv_id":"2502.08071","repositories_listed":1,"syntology":null},{"url":"/paper/generating-with-fairness-a-modality-diffused","title":"Generating with Fairness: A Modality-Diffused Counterfactual Framework for Incomplete Multimodal Recommendations","date":"2025-01-21","arxiv_id":"2501.11916","repositories_listed":1,"syntology":null},{"url":"/paper/dynamic-multimodal-fusion-via-meta-learning","title":"Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation","date":"2025-01-13","arxiv_id":"2501.07110","repositories_listed":1,"syntology":null},{"url":"/paper/spectrum-based-modality-representation-fusion","title":"Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recommendation","date":"2024-12-19","arxiv_id":"2412.14978","repositories_listed":1,"syntology":null},{"url":"/paper/modality-independent-graph-neural-networks","title":"Modality-Independent Graph Neural Networks with Global Transformers for Multimodal Recommendation","date":"2024-12-18","arxiv_id":"2412.13994","repositories_listed":1,"syntology":null},{"url":"/paper/beyond-graph-convolution-multimodal","title":"Beyond Graph Convolution: Multimodal Recommendation with Topology-aware MLPs","date":"2024-12-16","arxiv_id":"2412.11747","repositories_listed":1,"syntology":null},{"url":"/paper/stair-manipulating-collaborative-and","title":"STAIR: Manipulating Collaborative and Multimodal Information for E-Commerce Recommendation","date":"2024-12-16","arxiv_id":"2412.11729","repositories_listed":1,"syntology":null},{"url":"/paper/a-multimodal-single-branch-embedding-network","title":"A Multimodal Single-Branch Embedding Network for Recommendation in Cold-Start and Missing Modality Scenarios","date":"2024-09-26","arxiv_id":"2409.17864","repositories_listed":1,"syntology":null},{"url":"/paper/train-once-deploy-anywhere-matryoshka","title":"Train Once, Deploy Anywhere: Matryoshka Representation Learning for Multimodal Recommendation","date":"2024-09-25","arxiv_id":"2409.16627","repositories_listed":1,"syntology":null},{"url":"/paper/ducho-meets-elliot-large-scale-benchmarks-for-1","title":"Ducho meets Elliot: Large-scale Benchmarks for Multimodal Recommendation","date":"2024-09-24","arxiv_id":"2409.15857","repositories_listed":1,"syntology":null},{"url":"/paper/do-we-really-need-to-drop-items-with-missing","title":"Do We Really Need to Drop Items with Missing Modalities in Multimodal Recommendation?","date":"2024-08-21","arxiv_id":"2408.11767","repositories_listed":1,"syntology":{"n":5,"n_ran":4,"n_unverified":1,"n_pointer_only":5}},{"url":"/paper/harnessing-multimodal-large-language-models","title":"Harnessing Multimodal Large Language Models for Multimodal Sequential Recommendation","date":"2024-08-19","arxiv_id":"2408.09698","repositories_listed":1,"syntology":{"n":3,"n_ran":3,"n_unverified":0,"n_pointer_only":3}},{"url":"/paper/modality-balanced-learning-for-multimedia","title":"Modality-Balanced Learning for Multimedia Recommendation","date":"2024-07-26","arxiv_id":"2408.06360","repositories_listed":1,"syntology":null},{"url":"/paper/gume-graphs-and-user-modalities-enhancement","title":"GUME: Graphs and User Modalities Enhancement for Long-Tail Multimodal Recommendation","date":"2024-07-17","arxiv_id":"2407.12338","repositories_listed":1,"syntology":{"n":6,"n_ran":5,"n_unverified":1,"n_pointer_only":6}},{"url":"/paper/end-to-end-training-of-multimodal-model-and","title":"End-to-end training of Multimodal Model and ranking Model","date":"2024-04-09","arxiv_id":"2404.06078","repositories_listed":1,"syntology":null},{"url":"/paper/an-aligning-and-training-framework-for","title":"AlignRec: Aligning and Training in Multimodal Recommendations","date":"2024-03-19","arxiv_id":"2403.12384","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/ducho-2-0-towards-a-more-up-to-date-feature","title":"Ducho 2.0: Towards a More Up-to-Date Unified Framework for the Extraction of Multimodal Features in Recommendation","date":"2024-03-07","arxiv_id":"2403.04503","repositories_listed":1,"syntology":null},{"url":"/paper/mentor-multi-level-self-supervised-learning","title":"MENTOR: Multi-level Self-supervised Learning for Multimodal Recommendation","date":"2024-02-29","arxiv_id":"2402.19407","repositories_listed":1,"syntology":null},{"url":"/paper/disentangled-graph-variational-auto-encoder","title":"Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation with Interpretability","date":"2024-02-25","arxiv_id":"2402.16110","repositories_listed":1,"syntology":null},{"url":"/paper/mirror-gradient-towards-robust-multimodal","title":"Mirror Gradient: Towards Robust Multimodal Recommender Systems via Exploring Flat Local Minima","date":"2024-02-17","arxiv_id":"2402.11262","repositories_listed":1,"syntology":null},{"url":"/paper/lgmrec-local-and-global-graph-learning-for","title":"LGMRec: Local and Global Graph Learning for Multimodal Recommendation","date":"2023-12-27","arxiv_id":"2312.16400","repositories_listed":1,"syntology":null},{"url":"/paper/fmmrec-fairness-aware-multimodal","title":"Causality-Inspired Fair Representation Learning for Multimodal Recommendation","date":"2023-10-26","arxiv_id":"2310.17373","repositories_listed":1,"syntology":null},{"url":"/paper/semantic-guided-feature-distillation-for","title":"Semantic-Guided Feature Distillation for Multimodal Recommendation","date":"2023-08-06","arxiv_id":"2308.03113","repositories_listed":1,"syntology":null},{"url":"/paper/lightgt-a-light-graph-transformer-for","title":"LightGT: A Light Graph Transformer for Multimedia Recommendation","date":"2023-07-18","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/ducho-a-unified-framework-for-the-extraction","title":"Ducho: A Unified Framework for the Extraction of Multimodal Features in Recommendation","date":"2023-06-29","arxiv_id":"2306.17125","repositories_listed":1,"syntology":null},{"url":"/paper/mmrec-simplifying-multimodal-recommendation","title":"MMRec: Simplifying Multimodal Recommendation","date":"2023-02-02","arxiv_id":"2302.03497","repositories_listed":1,"syntology":null}],"syntology_records":4,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}