Papers › OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

1 Jan 2024CVPR 2024 1archive 2025-07-28

Siddharth Srivastava, Gaurav Sharma

We present a novel multimodal multitask network and associated training algorithm. The method is capable of ingesting data from approximately 12 different modalities namely image video audio text depth point cloud time series tabular graph X-ray infrared IMU and hyperspectral. The proposed approach utilizes modality specialized tokenizers a shared transformer architecture and cross-attention mechanisms to project the data from different modalities into a unified embedding space. It addresses multimodal and multitask scenarios by incorporating modality-specific task heads for different tasks in respective modalities. We propose a novel pretraining strategy with iterative modality switching to initialize the network and a training algorithm which trades off fully joint training over all modalities with training on pairs of modalities at a time. We provide comprehensive evaluation across 25 datasets from 12 modalities and show state of the art performances demonstrating the effectiveness of the proposed architecture pretraining strategy and adapted multitask training.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Point Cloud ClassificationAction ClassificationAction RecognitionAudio ClassificationFine-Grained Image ClassificationImage ClassificationSemantic SegmentationText SummarizationTime SeriesZero-Shot Video Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
3D Point Cloud Classification ModelNet40-C OmniVec2 Error Rate 0.142 #1 of 13 Archive leaderboard report
3D Point Cloud Classification ScanObjectNN OmniVec2 Overall Accuracy 97.2 #1 of 77 Archive leaderboard report
Action Classification Kinetics-400 OmniVec2 Acc@1 93.6 #1 of 207 Archive leaderboard report
Action Classification MiT OmniVec2 Top 1 Accuracy 53.1 #1 of 29 Archive leaderboard report
Action Classification Moments in Time OmniVec2 Top 1 Accuracy 53.1 #1 of 4 Archive leaderboard report
Action Recognition UCF101 OmniVec2 3-fold Accuracy 99.6 #4 of 91 Archive leaderboard report
Audio Classification AudioSet OmniVec2 Test mAP 0.558 #1 of 51 Archive leaderboard report
Audio Classification ESC-50 OmniVec2 Accuracy (5-fold) 99.1 #1 of 29 Archive leaderboard report
Audio Classification ESC-50 OmniVec2 PRE-TRAINING DATASET Multiple #1 of 29 Archive leaderboard report
Audio Classification ESC-50 OmniVec2 Top-1 Accuracy 99.1 #1 of 29 Archive leaderboard report
Fine-Grained Image Classification Oxford-IIIT Pet Dataset OmniVec2 Accuracy 99.6 #1 of 15 Archive leaderboard report
Image Classification ImageNet OmniVec2 Top 1 Accuracy 89.3% #22 of 1060 Archive leaderboard report
Image Classification Places365 OmniVec2 Top 1 Accuracy 65.1 #1 of 7 Archive leaderboard report
Image Classification iNaturalist 2018 OmniVec2 Top-1 Accuracy 94.6 #1 of 60 Archive leaderboard report
Semantic Segmentation NYU Depth v2 OmniVec2 Mean IoU 63.6 #1 of 121 Archive leaderboard report
Text Summarization DialogSum OmniVec2 BertScore 72.8 #2 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec2 Rouge1 47.6 #2 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec2 Rouge2 22.1 #2 of 4 Archive leaderboard report
Text Summarization DialogSum OmniVec2 RougeL 41.4 #2 of 4 Archive leaderboard report
Text Summarization SAMSum OmniVec2 BertScoreF1 65.1 #1 of 12 Archive leaderboard report
Text Summarization SAMSum OmniVec2 ROUGE-1 59.1 #1 of 12 Archive leaderboard report
Text Summarization SAMSum OmniVec2 ROUGE-2 34.1 #1 of 12 Archive leaderboard report
Text Summarization SAMSum OmniVec2 ROUGE-L 63.7 #1 of 12 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 OmniVec2 text-to-video R@1 26.1 #1 of 9 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 OmniVec2 text-to-video R@10 70.8 #1 of 9 Archive leaderboard report
Zero-Shot Video Retrieval YouCook2 OmniVec2 text-to-video R@5 54.1 #1 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

I3DR-Net

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections