Papers › OmniVec: Learning robust representations with cross modal sharing

OmniVec: Learning robust representations with cross modal sharing

7 Nov 2023arXiv:2311.05709archive 2025-07-28

Siddharth Srivastava, Gaurav Sharma

Majority of research in learning based methods has been towards designing and training networks for specific tasks. However, many of the learning based tasks, across modalities, share commonalities and could be potentially tackled in a joint framework. We present an approach in such direction, to learn multiple tasks, in multiple modalities, with a unified architecture. The proposed network is composed of task specific encoders, a common trunk in the middle, followed by task specific prediction heads. We first pre-train it by self-supervised masked training, followed by sequential training for the different tasks. We train the network on all major modalities, e.g.\ visual, audio, text and 3D, and report results on $22$ diverse and challenging public benchmarks. We demonstrate empirically that, using a joint network to train across modalities leads to meaningful information sharing and this allows us to achieve state-of-the-art results on most of the benchmarks. We also show generalization of the trained network on cross-modal tasks as well as unseen datasets and tasks.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Point Cloud ClassificationAction ClassificationAudio ClassificationFine-Grained Image ClassificationImage ClassificationSemantic SegmentationText Summarization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Point Cloud Classification ModelNet40-C OmniVec Error Rate 0.156 #2 of 13 Archive leaderboard report
3D Point Cloud Classification ScanObjectNN OmniVec Overall Accuracy 96.1 #3 of 77 Archive leaderboard report
Action Classification Kinetics-400 OmniVec Acc@1 91.1 #6 of 207 Archive leaderboard report
Action Classification MIT OmniVec Top 1 Accuracy 49.8 #2 of 2 Archive leaderboard report
Action Classification Moments in Time OmniVec Top 1 Accuracy 49.8 #2 of 4 Archive leaderboard report
Action Recognition UCF101 OmniVec 3-fold Accuracy 99.6 #3 of 91 Archive leaderboard report
Audio Classification AudioSet OmniVec Test mAP 0.548 #2 of 51 Archive leaderboard report
Audio Classification ESC-50 OmniVec Accuracy (5-fold) 98.4 #4 of 29 Archive leaderboard report
Audio Classification ESC-50 OmniVec PRE-TRAINING DATASET Multiple #4 of 29 Archive leaderboard report
Audio Classification ESC-50 OmniVec Top-1 Accuracy 98.4 #4 of 29 Archive leaderboard report
Fine-Grained Image Classification Oxford-IIIT Pet Dataset OmniVec Accuracy 99.2 #2 of 15 Archive leaderboard report
Image Classification Places365 OmniVec(ViT) Top 1 Accuracy 63.5 #2 of 7 Archive leaderboard report
Image Classification iNaturalist 2018 OmniVec Top-1 Accuracy 93.8 #2 of 60 Archive leaderboard report
Semantic Segmentation NYU Depth v2 OmniVec Mean IoU 60.8 #5 of 121 Archive leaderboard report
Semantic Segmentation S3DIS Area5 OmniVec mIoU 75.9 #2 of 61 Archive leaderboard report
Text Summarization DialogSum OmniVec BertScore 71.91 #3 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec Rouge1 46.91 #3 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec Rouge2 21.22 #3 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec RougeL 40.19 #3 of 4 Archive leaderboard report
Video Retrieval MSR-VTT-1kA OmniVec text-to-video R@10 89.4 #60 of 63 Archive leaderboard report
Video Retrieval MSR-VTT-1kA OmniVec (pretrained) text-to-video R@10 78.6 #62 of 63 Archive leaderboard report
Video Retrieval YouCook2 OmniVec text-to-video R@10 70.8 #15 of 16 Archive leaderboard report
Video Retrieval YouCook2 OmniVec (pretrained) text-to-video R@10 64.2 #16 of 16 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections