Browse State-of-the-Art › Grounded language learning
Grounded language learning
23 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Acquire the meaning of language in situated environments.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
23 shown of 23 papers with code (56 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
18 Oct 2018 6 repositories listed Syntology ran 0 of 14 samples · 14 unverified · 1 pointer-only (licence)Allowing humans to interactively train artificial agents to understand language instructions is desirable for both practical and scientific reasons, but given the poor data efficiency of the current learning methods,…
-
21 Mar 2024 2 repositories listedToday's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human…
-
3 Sep 2020 2 repositories listedRecent work has shown that large text-based neural language models, trained with conventional supervised learning objectives, acquire a surprising propensity for few- and one-shot learning.
-
2 May 2020 2 repositories listedTo study this human-like language acquisition ability, we present VisCOLL, a visually grounded language learning task, which simulates the continual acquisition of compositional phrases from streaming visual scenes.
-
2 Apr 2019 2 repositories listedWe improve the informativeness of models for conditional text generation using techniques from computational pragmatics.
-
6 Jul 2022 1 repository listed Syntology ran 7 of 10 samples · 3 unverifiedWe provide a study of how induced model sparsity can help achieve compositional generalization and better sample efficiency in grounded language learning problems.
-
5 Jul 2022 1 repository listedLexical semantics and cognitive science point to affordances (i.
-
22 Feb 2022 1 repository listedAfter training on an augmented dataset with almost forty times more adverbs than the original problem, a non-modular baseline is not able to systematically generalize to a novel combination of a known verb and adverb.
-
21 Feb 2022 1 repository listedIn this paper we create visually grounded word embeddings by combining English text and images and compare them to popular text-based methods, to see if visual information allows our model to better capture cognitive…
-
20 Oct 2021 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)We hope SILG enables the community to quickly identify new methodologies for language grounding that generalize to a diverse set of environments and their associated challenges.
-
16 Jun 2021 1 repository listedThis study addresses the question whether visually grounded speech recognition (VGS) models learn to capture sentence semantics without access to any prior linguistic knowledge.
-
13 Feb 2021 1 repository listedWe present a novel interactive learning protocol that enables training request-fulfilling agents by verbally describing their activities.
-
15 Apr 2020 1 repository listedTo facilitate the research on language-guided agents with domain adaption, we propose a novel zero-shot compositional policy learning task, where the environments are characterized as a composition of different…
-
4 Nov 2019 1 repository listedAlthough their encodeing method is not compositional like natural languages from a perspective of human beings, the emergent languages can be generalised to unseen inputs and, more importantly, are easier for models to…
-
9 Sep 2019 1 repository listedHumans learn language by interaction with their environment and listening to other humans.
-
Learning semantic sentence representations from visually grounded language without lexical knowledge27 Mar 2019 1 repository listedThe system achieves state-of-the-art results on several of these benchmarks, which shows that a system trained solely on multimodal data, without assuming any word representations, is able to capture sentence level…
-
26 Nov 2018 1 repository listedWe introduce a new inference task - Visual Entailment (VE) - which differs from traditional Textual Entailment (TE) tasks whereby a premise is defined by an image, rather than a natural language sentence as in TE tasks.
-
1 Oct 2018 1 repository listedPrevious work on grounded language learning did not fully capture the semantics underlying the correspondences between structured world state representations and texts, especially those between numerical values and…
-
20 Sep 2018 1 repository listedRecent work has shown how to learn better visual-semantic embeddings by leveraging image descriptions in more than one language.
-
21 May 2018 1 repository listedOur goal is to develop a model that can learn to follow new instructions given prior instruction-perception-action examples.
-
Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game26 Apr 2018 1 repository listedBuilding intelligent agents that can communicate with and learn from humans in natural language is of great value.
-
20 Jun 2017 1 repository listedTrained via a combination of reinforcement and unsupervised learning, and beginning with minimal prior knowledge, the agent learns to relate linguistic symbols to emergent perceptual representations of its physical…
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections