{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-sentence-embeddings-from-pre-trained","title":"On the Sentence Embeddings from Pre-trained Language Models","arxiv_id":"2011.05864","date":"2020-11-02","proceeding":"EMNLP 2020 11","authors":["Bohan Li","Hao Zhou","Junxian He","Mingxuan Wang","Yiming Yang","Lei LI"],"abstract":"Pre-trained contextual representations like BERT have achieved great success in natural language processing. However, the sentence embeddings from the pre-trained language models without fine-tuning have been found to poorly capture semantic meaning of sentences. In this paper, we argue that the semantic information in the BERT embeddings is not fully exploited. We first reveal the theoretical connection between the masked language model pre-training objective and the semantic similarity task theoretically, and then analyze the BERT sentence embeddings empirically. We find that BERT always induces a non-smooth anisotropic semantic space of sentences, which harms its performance of semantic similarity. To address this issue, we propose to transform the anisotropic sentence embedding distribution to a smooth and isotropic Gaussian distribution through normalizing flows that are learned with an unsupervised objective. Experimental results show that our proposed BERT-flow method obtains significant performance gains over the state-of-the-art sentence embeddings on a variety of semantic textual similarity tasks. The code is available at https://github.com/bohanli/BERT-flow.","url_abs":"https://arxiv.org/abs/2011.05864v1","url_pdf":"https://arxiv.org/pdf/2011.05864v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-sentence-embeddings-from-pre-trained","repo_url":"https://github.com/bohanli/BERT-flow","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"on-the-sentence-embeddings-from-pre-trained","repo_url":"https://github.com/InsaneLife/dssm","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}},{"paper_slug":"on-the-sentence-embeddings-from-pre-trained","repo_url":"https://github.com/sleepthroughdifficulties/kernelwhitening","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-embedding","task_name":"Sentence Embedding"},{"task_slug":"sentence-embeddings","task_name":"Sentence Embeddings"},{"task_slug":"sentence-embedding-1","task_name":"Sentence-Embedding"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bert","method_name":"BERT"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-linear-decay","method_name":"Linear Warmup With Linear Decay"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"normalizing-flows","method_name":"Normalizing Flows"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"weight-decay","method_name":"Weight Decay"},{"method_slug":"wordpiece","method_name":"WordPiece"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semantic-textual-similarity-on-sick","task":"Semantic Textual Similarity","dataset":"SICK","model":"BERTbase-flow (NLI)","rank_in_archive_order":21,"of":22,"metrics":{"Spearman Correlation":"0.6544"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts-benchmark","task":"Semantic Textual Similarity","dataset":"STS Benchmark","model":"BERTlarge-flow (target)","rank_in_archive_order":60,"of":66,"metrics":{"Spearman Correlation":"0.7226"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts12","task":"Semantic Textual Similarity","dataset":"STS12","model":"BERTlarge-flow (target)","rank_in_archive_order":18,"of":20,"metrics":{"Spearman Correlation":"0.6520"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts13","task":"Semantic Textual Similarity","dataset":"STS13","model":"BERTlarge-flow (target)","rank_in_archive_order":21,"of":22,"metrics":{"Spearman Correlation":"0.7339"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts14","task":"Semantic Textual Similarity","dataset":"STS14","model":"BERTlarge-flow (target)","rank_in_archive_order":20,"of":21,"metrics":{"Spearman Correlation":"0.6942"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts15","task":"Semantic Textual Similarity","dataset":"STS15","model":"BERTlarge-flow (target)","rank_in_archive_order":20,"of":20,"metrics":{"Spearman Correlation":"0.7492"},"uses_additional_data":false},{"leaderboard":"/sota/semantic-textual-similarity-on-sts16","task":"Semantic Textual Similarity","dataset":"STS16","model":"BERTlarge-flow (target)","rank_in_archive_order":16,"of":20,"metrics":{"Spearman Correlation":"0.7763"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2011.05864","atlas_url":"https://app.syntology.ai/?focus=2011.05864","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2011.05864"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sleepthroughdifficulties/kernelwhitening","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bohanli/BERT-flow","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/InsaneLife/dssm","reach":{"status":"unanswered"}}],"summary":{"ran_honours":2,"ran_violates":1},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b8d3902cb6058201","entry":"get_gt","repo":"bohanli/BERT-flow","repo_kind":"official","path":"scripts/eval_stsb.py","file_url":"https://github.com/bohanli/BERT-flow/blob/HEAD/scripts/eval_stsb.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b8d3902cb6058201"}},{"code_sha256_prefix":"3b6418910c030c74","entry":"get_pred","repo":"bohanli/BERT-flow","repo_kind":"official","path":"scripts/eval_stsb.py","file_url":"https://github.com/bohanli/BERT-flow/blob/HEAD/scripts/eval_stsb.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"3b6418910c030c74"}},{"code_sha256_prefix":"c801efb328b1fc65","entry":"pearson_r","repo":"bohanli/BERT-flow","repo_kind":"official","path":"scripts/eval_stsb.py","file_url":"https://github.com/bohanli/BERT-flow/blob/HEAD/scripts/eval_stsb.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c801efb328b1fc65"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}