Browse State-of-the-Art › Language Acquisition
Language Acquisition
83 papers with code · 1 benchmark · 7 datasets archive 2025-07-28
Language acquisition refers to tasks related to the learning of a second language.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SLAM 2018 (1 row) | Context Based Model | Context Based Approach for Second Language Acquisition | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 83 papers with code (522 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Mar 2024 2 repositories listedToday's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human…
-
20 Oct 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)But to achieve these results, LMs must be trained in distinctly un-human-like ways - requiring orders of magnitude more language data than children receive during development, and without perceptual or social context.
-
22 Sep 2021 2 repositories listedWe then demonstrate the utility of the compiled corpora through (1) a longitudinal corpus study of the prevalence of different syntactic and semantic phenomena in the CDS, and (2) applying an existing computational…
-
2 May 2020 2 repositories listedTo study this human-like language acquisition ability, we present VisCOLL, a visually grounded language learning task, which simulates the continual acquisition of compositional phrases from streaming visual scenes.
-
30 Apr 2020 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We propose transfer learning as a method for analyzing the encoding of grammatical structure in neural language models.
-
21 Feb 2020 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)We argue that the ability to imagine out-of-distribution goals is key to enable creative discoveries and open-ended learning.
-
25 Aug 2019 2 repositories listedSecond language acquisition (SLA) modeling is to predict whether second language learners could correctly answer the questions according to what they have learned.
-
1 Jun 2018 2 repositories listedOur system uses a logistic regression model to predict the likelihood of a student making a mistake while answering an exercise on Duolingo in all three language tracks - English/Spanish (en/es), Spanish/English (es/en)…
-
31 May 2018 2 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedThis paper investigates the ability of artificial neural networks to judge the grammatical acceptability of a sentence, with the goal of testing their linguistic competence.
-
31 Jan 2018 2 repositories listedWe build a virtual agent for learning language in a 2D maze-like world.
-
29 Nov 2017 2 repositories listedRecent work has attempted to characterize the structure of semantic memory and the search algorithms which, together, best approximate human patterns of search revealed in a semantic fluency task.
-
5 Oct 2017 2 repositories listedWe introduce a newly collected data set of human semantic relevance judgements and an associated task, semantic speech retrieval, where the goal is to search for spoken utterances that are semantically relevant to a…
-
17 Feb 2025 1 repository listedExperiment results on CEFR-SP and TurkCorpus datasets show that the proposed method can effectively increase the frequency and diversity of vocabulary of the target level by more than 20% compared to baseline models,…
-
11 Feb 2025 1 repository listedHowever, they are hard to control for linguistic forms that correspond to learners' current needs, such as grammar.
-
9 Nov 2024 1 repository listedWhether and how language models (LMs) acquire the syntax of natural languages has been widely evaluated under the minimal pair paradigm.
-
30 Oct 2024 1 repository listedCurriculum Learning has been a popular strategy to improve the cognitive plausibility of Small-Scale Language Models (SSLMs) in the BabyLM Challenge.
-
30 Oct 2024 1 repository listedWe apply this pipeline to the 100-million-word pre-training dataset from the BabyLM challenge, as well as to standard language and grammatical benchmarks, enabling us to pre-train and evaluate a model using phonemic…
-
23 Oct 2024 1 repository listedHumans develop their grammars by making structural generalizations from finite input.
-
17 Oct 2024 1 repository listedNotably, in generation tasks, LMs are more similar to human performance in areas where information is easier to extract from the corpus, such as average word length, clauses, and auxiliary verbs.
-
9 Aug 2024 1 repository listedThis gives rise to a novel hypothesis that CDG is facilitated insofar as the features of the exposure context--in particular, its first postverbal argument--are harmonically aligned.
-
27 Jul 2024 1 repository listedWe vary the age of exposure by training LMs on language pairs in various experimental conditions, and find that LMs, which lack any direct analog to innate maturational stages, do not show CP effects when the age of…
-
19 Jun 2024 1 repository listed Syntology ran 14 of 14 samples · 0 unverified · 14 pointer-only (licence)Recent studies have compared the analogical reasoning abilities of human subjects and Large Language Models (LLMs) on abstract symbol manipulation tasks, such as letter string analogies.
-
17 Jun 2024 1 repository listedHere, we argue that they are instead both contingent on a more general learning strategy for language acquisition: joint learning.
-
22 May 2024 1 repository listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)In this work, we explore how corrective feedback from interactions influences neural language acquisition from scratch through systematically controlled experiments, assessing whether it contributes to word learning…
-
13 May 2024 1 repository listedChild-directed speech (CDS) is a particular type of speech that adults use when addressing young children.
-
18 Feb 2024 1 repository listedHowever, it is unclear whether or how these models represent grammatical information from the learned languages.
-
7 Nov 2023 1 repository listedIn this work, we create MMA, a large, flexible, multilingual, and multi-domain dataset of informal-formal pairs, by using a language model to translate in the reverse direction, that is, from formal mathematical…
-
6 Nov 2023 1 repository listedIn this paper, we describe the University of Lyon 2 submission to the Strict-Small track of the BabyLM competition.
-
31 Oct 2023 1 repository listedThe success of neural language models (LMs) on many technological tasks has brought about their potential relevance as scientific theories of language despite some clear differences between LM training and child…
-
31 Oct 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedWe employ a characterization of linguistic complexity from psycholinguistic and language acquisition research to develop data-driven curricula to understand the underlying linguistic knowledge that models learn to…
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections