Papers › Text Embeddings by Weakly-Supervised Contrastive Pre-training

Text Embeddings by Weakly-Supervised Contrastive Pre-training

7 Dec 2022arXiv:2212.03533archive 2025-07-28

Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, Furu Wei

This paper presents E5, a family of state-of-the-art text embeddings that transfer well to a wide range of tasks. The model is trained in a contrastive manner with weak supervision signals from our curated large-scale text pair dataset (called CCPairs). E5 can be readily used as a general-purpose embedding model for any tasks requiring a single-vector representation of texts such as retrieval, clustering, and classification, achieving strong performance in both zero-shot and fine-tuned settings. We conduct extensive evaluations on 56 datasets from the BEIR and MTEB benchmarks. For zero-shot settings, E5 is the first model that outperforms the strong BM25 baseline on the BEIR retrieval benchmark without using any labeled data. When fine-tuned, E5 obtains the best results on the MTEB benchmark, beating existing embedding models with 40x more parameters.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

microsoft/unilm officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

MTEB BenchmarkOnly Connect Walls Dataset Task 1 (Grouping)Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) Wasserstein Distance (WD) 83.8 ± .6 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) # Correct Groups 89 ± 6 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) # Solved Walls 1 ± 0 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) Adjusted Mutual Information (AMI) 19.5 ± .4 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) Adjusted Rand Index (ARI) 16.3 ± .4 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (BASE) Fowlkes Mallows Score (FMS) 33.1 ± .3 #11 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) Wasserstein Distance (WD) 84.4 ± .7 #13 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) # Correct Groups 76 ± 5 #13 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) # Solved Walls 0 ± 0 #13 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) Adjusted Mutual Information (AMI) 18.5 ± .6 #13 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) Adjusted Rand Index (ARI) 15.4 ± .5 #13 of 22 Archive leaderboard report
Only Connect Walls Dataset Task 1 (Grouping) OCW E5 (LARGE) Fowlkes Mallows Score (FMS) 32.3 ± .4 #13 of 22 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections