Papers › VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on...

VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

14 Dec 2021ACL 2022 5arXiv:2112.07566archive 2025-07-28

Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan, Anette Frank, Iacer Calixto, Albert Gatt

We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2112.07566")

Code

Syntology Ran 4 of 9 code samples harvested from 1 repository linked to this paper; 5 have no recorded run. Of those that ran: 4 ran with no contract checked.

By repository: official repository: 9 samples from 1 repository, 4 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

heidelberg-nlp/valse officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

9 samples harvested; 4 ran; 0 honoured the contract we drafted; 5 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

4ran
5unverified

Licence: 0 of the 9 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from heidelberg-nlp/valse. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

do_nms heidelberg-nlp/valse/modeling_frcnn.py official repository ran MIT (permissive) · e21bb85b8e1b4836 · report
is_remote_url heidelberg-nlp/valse/utils.py official repository ran MIT (permissive) · 899d1e2736236351 · report
load_labels heidelberg-nlp/valse/utils.py official repository ran MIT (permissive) · 8a78a755bda274a5 · report
norm_box heidelberg-nlp/valse/modeling_frcnn.py official repository ran fingerprinted MIT (permissive) · cfa4887da544c990 · report
compute_ppl heidelberg-nlp/valse/unimodal_valse_eval.py official repository unverified MIT (permissive) · 3b772f0bbcecca2c · report
load_checkpoint heidelberg-nlp/valse/utils.py official repository unverified MIT (permissive) · bd3a3f4bd8750da2 · report
pad_list_tensors heidelberg-nlp/valse/modeling_frcnn.py official repository unverified MIT (permissive) · c06bd30b6fcc7498 · report
read_foil_dataset heidelberg-nlp/valse/read_foil_dataset.py official repository unverified MIT (permissive) · 0ea1a5e4fb771a67 · report
read_foils heidelberg-nlp/valse/read_foil_dataset.py official repository unverified MIT (permissive) · dc379b20add70eab · report

Tasks

image-sentence alignment

1 archive task tag without a task page not shown.

Datasets

Introduced by this paper, per the archive.

VALSE

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
image-sentence alignment VALSE ViLBERT 12-in-1 Average Accuracy 63.2 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE ViLBERT 12-in-1 average pairwise accuracy 75.1 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE CLIP average pairwise accuracy 64.0 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE ViLBERT Average Accuracy 51.3 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE ViLBERT average pairwise accuracy 63.7 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE GPT1 average pairwise accuracy 60.7 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE GPT2 average pairwise accuracy 60.1 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE LXMERT Average Accuracy 53.5 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE LXMERT average pairwise accuracy 59.6 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE VisualBERT Average Accuracy 48.8 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE VisualBERT average pairwise accuracy 46.4 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap GPT2 pairwise accuracy 76.9 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap GPT1 pairwise accuracy 72.2 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap CLIP pairwise accuracy 68.6 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap ViLBERT Accuracy (%) 50.4 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap ViLBERT pairwise accuracy 68.3 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap ViLBERT 12-in-1 Accuracy (%) 52.2 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap ViLBERT 12-in-1 pairwise accuracy 58.9 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap LXMERT Accuracy (%) 48.5 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap LXMERT pairwise accuracy 45.8 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap VisualBERT Accuracy (%) 49.7 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE actant swap VisualBERT pairwise accuracy 44.4 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement CLIP pairwise accuracy 75.6 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement ViLBERT Accuracy (%) 52.6 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement ViLBERT pairwise accuracy 70.7 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement GPT2 pairwise accuracy 66.8 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement ViLBERT 12-in-1 Accuracy (%) 57.3 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement ViLBERT 12-in-1 pairwise accuracy 65.9 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement GPT1 pairwise accuracy 65.4 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement LXMERT Accuracy (%) 51.1 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement LXMERT pairwise accuracy 54.8 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement VisualBERT Accuracy (%) 48.8 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE action replacement VisualBERT pairwise accuracy 49.2 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean ViLBERT 12-in-1 Accuracy (%) 54.3 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean ViLBERT 12-in-1 pairwise accuracy 69.2 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean GPT2 pairwise accuracy 50.0 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean CLIP pairwise accuracy 49.7 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean ViLBERT Accuracy (%) 50.0 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean ViLBERT pairwise accuracy 48.1 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean VisualBERT Accuracy (%) 50.0 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean VisualBERT pairwise accuracy 47.6 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean GPT1 pairwise accuracy 45.2 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean LXMERT Accuracy (%) 49.0 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference clean LXMERT pairwise accuracy 44.2 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard ViLBERT 12-in-1 Accuracy (%) 54.4 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard ViLBERT 12-in-1 pairwise accuracy 75.7 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard GPT2 pairwise accuracy 54.5 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard CLIP pairwise accuracy 52.1 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard VisualBERT Accuracy (%) 50.0 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard VisualBERT pairwise accuracy 49.5 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard ViLBERT Accuracy (%) 50.0 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard ViLBERT pairwise accuracy 47.2 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard LXMERT Accuracy (%) 49.8 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard LXMERT pairwise accuracy 46.8 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE coreference standard GPT1 pairwise accuracy 45.6 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial ViLBERT 12-in-1 Accuracy (%) 66.7 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial ViLBERT 12-in-1 pairwise accuracy 77.3 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial ViLBERT Accuracy (%) 51.8 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial ViLBERT pairwise accuracy 73.7 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial GPT1 pairwise accuracy 69.5 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial CLIP pairwise accuracy 57.5 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial VisualBERT Accuracy (%) 50.0 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial VisualBERT pairwise accuracy 50.0 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial GPT2 pairwise accuracy 45.3 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial LXMERT Accuracy (%) 49.9 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting adversarial LXMERT pairwise accuracy 42.6 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced ViLBERT 12-in-1 Accuracy (%) 64.9 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced ViLBERT 12-in-1 pairwise accuracy 76.7 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced LXMERT Accuracy (%) 52.0 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced LXMERT pairwise accuracy 62.2 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced CLIP pairwise accuracy 62.1 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced ViLBERT Accuracy (%) 50.7 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced ViLBERT pairwise accuracy 58.6 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced GPT2 pairwise accuracy 51.6 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced GPT1 pairwise accuracy 51.2 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced VisualBERT Accuracy (%) 48.3 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting balanced VisualBERT pairwise accuracy 48.2 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers ViLBERT 12-in-1 Accuracy (%) 69.2 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers ViLBERT 12-in-1 pairwise accuracy 80.2 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers LXMERT Accuracy (%) 55.4 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers LXMERT pairwise accuracy 69.2 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers ViLBERT Accuracy (%) 50.6 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers ViLBERT pairwise accuracy 62.9 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers CLIP pairwise accuracy 62.5 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers GPT2 pairwise accuracy 49.8 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers GPT1 pairwise accuracy 48.7 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers VisualBERT Accuracy (%) 47.8 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE counting small numbers VisualBERT pairwise accuracy 48.2 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE existence ViLBERT 12-in-1 Accuracy (%) 89.0 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE existence ViLBERT 12-in-1 pairwise accuracy 95.6 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE existence LXMERT Accuracy (%) 55.8 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE existence LXMERT pairwise accuracy 78.6 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE existence CLIP pairwise accuracy 66.9 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE existence ViLBERT Accuracy (%) 2.4 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE existence ViLBERT pairwise accuracy 66.5 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE existence GPT1 pairwise accuracy 61.8 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE existence GPT2 pairwise accuracy 58.0 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE existence VisualBERT Accuracy (%) 49.3 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE existence VisualBERT pairwise accuracy 39.7 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) CLIP pairwise accuracy 88.8 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) LXMERT Accuracy (%) 70.8 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) LXMERT pairwise accuracy 87.1 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) ViLBERT 12-in-1 Accuracy (%) 71.5 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) ViLBERT 12-in-1 pairwise accuracy 86.9 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) ViLBERT Accuracy (%) 55.9 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) ViLBERT pairwise accuracy 86.9 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) GPT2 pairwise accuracy 80.7 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) GPT1 pairwise accuracy 77.5 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) VisualBERT Accuracy (%) 46.6 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE foil-it (noun phrases) VisualBERT pairwise accuracy 48.5 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality ViLBERT 12-in-1 Accuracy (%) 62.0 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality ViLBERT 12-in-1 pairwise accuracy 72.4 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality LXMERT Accuracy (%) 55.1 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality LXMERT pairwise accuracy 64.4 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality ViLBERT Accuracy (%) 50.3 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality ViLBERT pairwise accuracy 61.2 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality CLIP pairwise accuracy 56.2 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality GPT1 pairwise accuracy 53.1 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality GPT2 pairwise accuracy 51.9 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality VisualBERT Accuracy (%) 46.5 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE plurality VisualBERT pairwise accuracy 45.7 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations GPT1 pairwise accuracy 77.2 #1 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations GPT2 pairwise accuracy 75.0 #2 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations ViLBERT 12-in-1 Accuracy (%) 53.4 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations ViLBERT 12-in-1 pairwise accuracy 67.7 #3 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations CLIP pairwise accuracy 64.3 #4 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations LXMERT Accuracy (%) 50.8 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations LXMERT pairwise accuracy 60.2 #5 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations ViLBERT Accuracy (%) 49.9 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations ViLBERT pairwise accuracy 57.2 #6 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations VisualBERT Accuracy (%) 49.3 #7 of 7 Archive leaderboard report
image-sentence alignment VALSE spatial relations VisualBERT pairwise accuracy 39.7 #7 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections