Papers › Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis

Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis

11 Feb 2023arXiv:2302.05608archive 2025-07-28

Zhu Wang, Sourav Medya, Sathya N. Ravi

Often, deep network models are purely inductive during training and while performing inference on unseen data. Thus, when such models are used for predictions, it is well known that they often fail to capture the semantic information and implicit dependencies that exist among objects (or concepts) on a population level. Moreover, it is still unclear how domain or prior modal knowledge can be specified in a backpropagation friendly manner, especially in large-scale and noisy settings. In this work, we propose an end-to-end vision and language model incorporating explicit knowledge graphs. We also introduce an interactive out-of-distribution (OOD) layer using implicit network operator. The layer is used to filter noise that is brought by external knowledge base. In practice, we apply our model on several vision and language downstream tasks including visual question answering, visual reasoning, and image-text retrieval on different datasets. Our experiments show that it is possible to design models that perform similarly to state-of-art results but with significantly fewer samples and training time.

PaperPDFCode

Code

ellenzhuwang/VK_OOD officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image-text RetrievalKnowledge GraphsLanguage ModelingLanguage ModellingOutlier DetectionQuestion AnsweringRetrievalText RetrievalVisual Question AnsweringVisual Question Answering (VQA)Visual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Question Answering VQA v2 test-dev VK-OOD Accuracy 76.8 #9 of 11 Archive leaderboard report
Visual Question Answering (VQA) OK-VQA VK-OOD Accuracy 52.4 #15 of 37 Archive leaderboard report
Visual Reasoning NLVR2 Dev VK-OOD Accuracy 83.9 #10 of 15 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPECLIPDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformerfail

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections