Papers › MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion...

MixMAS: A Framework for Sampling-Based Mixer Architecture Search for Multimodal Fusion and Learning

24 Dec 2024arXiv:2412.18437archive 2025-07-28

Abdelmadjid Chergui, Grigor Bezirganyan, Sana Sellami, Laure Berti-ÉQuille, Sébastien Fournier

Choosing a suitable deep learning architecture for multimodal data fusion is a challenging task, as it requires the effective integration and processing of diverse data types, each with distinct structures and characteristics. In this paper, we introduce MixMAS, a novel framework for sampling-based mixer architecture search tailored to multimodal learning. Our approach automatically selects the optimal MLP-based architecture for a given multimodal machine learning (MML) task. Specifically, MixMAS utilizes a sampling-based micro-benchmarking strategy to explore various combinations of modality-specific encoders, fusion functions, and fusion networks, systematically identifying the architecture that best meets the task's performance metrics.

PaperPDFCode

Code

Madjid-CH/auto-mixer officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Benchmarking

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections