Papers › Not all layers are equally as important: Every Layer Counts BERT

Not all layers are equally as important: Every Layer Counts BERT

3 Nov 2023arXiv:2311.02265archive 2025-07-28

Lucas Georges Gabriel Charpentier, David Samuel

This paper introduces a novel modification of the transformer architecture, tailored for the data-efficient pretraining of language models. This aspect is evaluated by participating in the BabyLM challenge, where our solution won both the strict and strict-small tracks. Our approach allows each transformer layer to select which outputs of previous layers to process. The empirical results verify the potential of this simple modification and show that not all layers are equally as important.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AllLinguistic AcceptabilityNatural Language Inference

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Linguistic Acceptability CoLA LTG-BERT-base 98M Accuracy 82.7 #6 of 43 Archive leaderboard report
Linguistic Acceptability CoLA ELC-BERT-base 98M Accuracy 82.6 #7 of 43 Archive leaderboard report
Linguistic Acceptability CoLA LTG-BERT-small 24M Accuracy 77.6 #10 of 43 Archive leaderboard report
Linguistic Acceptability CoLA ELC-BERT-small 24M Accuracy 76.1 #11 of 43 Archive leaderboard report
Natural Language Inference MultiNLI ELC-BERT-base 98M (zero init) Matched 84.4 #32 of 67 Archive leaderboard report
Natural Language Inference MultiNLI ELC-BERT-base 98M (zero init) Mismatched 84.5 #32 of 67 Archive leaderboard report
Natural Language Inference MultiNLI LTG-BERT-base 98M Matched 83 #36 of 67 Archive leaderboard report
Natural Language Inference MultiNLI LTG-BERT-base 98M Mismatched 83.4 #36 of 67 Archive leaderboard report
Natural Language Inference MultiNLI ELC-BERT-small 24M Matched 79.2 #44 of 67 Archive leaderboard report
Natural Language Inference MultiNLI ELC-BERT-small 24M Mismatched 79.9 #44 of 67 Archive leaderboard report
Natural Language Inference MultiNLI LTG-BERT-small 24M Matched 78 #45 of 67 Archive leaderboard report
Natural Language Inference MultiNLI LTG-BERT-small 24M Mismatched 78.8 #45 of 67 Archive leaderboard report
Natural Language Inference RTE ELC-BERT-base 98M (zero init) Accuracy 63 #67 of 90 Archive leaderboard report
Natural Language Inference RTE ELC-BERT-small 24M Accuracy 55.4 #82 of 90 Archive leaderboard report
Natural Language Inference RTE LTG-BERT-base 98M Accuracy 54.7 #84 of 90 Archive leaderboard report
Natural Language Inference RTE LTG-BERT-small 24M Accuracy 53.7 #87 of 90 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections