Papers › Deep Multimodal Subspace Clustering Networks
Deep Multimodal Subspace Clustering Networks
Mahdi Abavisani, Vishal M. Patel
We present convolutional neural network (CNN) based approaches for unsupervised multimodal subspace clustering. The proposed framework consists of three main stages - multimodal encoder, self-expressive layer, and multimodal decoder. The encoder takes multimodal data as input and fuses them to a latent space representation. The self-expressive layer is responsible for enforcing the self-expressiveness property and acquiring an affinity matrix corresponding to the data points. The decoder reconstructs the original input data. The network uses the distance between the decoder's reconstruction and the original input in its training. We investigate early, late and intermediate fusion techniques and propose three different encoders corresponding to them for spatial fusion. The self-expressive layers and multimodal decoders are essentially the same for different spatial fusion-based approaches. In addition to various spatial fusion-based methods, an affinity fusion-based network is also proposed in which the self-expressive layer corresponding to different modalities is enforced to be the same. Extensive experiments on three datasets show that the proposed methods significantly outperform the state-of-the-art multimodal subspace clustering methods.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Image Clustering | ARL Polarimetric Thermal Face Dataset | DMSC | Accuracy | 0.983 | #1 of 1 | Archive leaderboard | report |
| Image Clustering | Extended Yale-B | DMSC | Accuracy | 0.992 | #1 of 9 | Archive leaderboard | report |
| Image Clustering | Extended Yale-B | DMSC | NMI | 0.988 | #1 of 9 | Archive leaderboard | report |
| Image Clustering | USPS | DMSC | Accuracy | 0.951 | #7 of 16 | Archive leaderboard | report |
| Image Clustering | USPS | DMSC | NMI | 0.929 | #7 of 16 | Archive leaderboard | report |
| Multi-view Subspace Clustering | ARL Polarimetric Thermal Face Dataset | DMSC | Accuracy | 0.988 | #1 of 1 | Archive leaderboard | report |
| Multi-view Subspace Clustering | ORL | DMSC | Accuracy | 0.833 | #2 of 3 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections