{"url":"/dataset/discovery","name":"Discovery","full_name":null,"description_markdown":"The *Discovery* datasets consists of adjacent sentence pairs (s1,s2) with a discourse marker (y) that occurred at the beginning of s2. They were extracted from the depcc web corpus.\r\n\r\nMarkers prediction can be used in order to train a sentence encoders. Discourse markers can be considered as noisy labels for various semantic tasks, such as entailment (y=therefore), subjectivity analysis (y=personally) or sentiment analysis (y=sadly), similarity (y=similarly), typicality, (y=curiously) ...\r\n\r\nThe specificity of this dataset is the diversity of the markers, since previously used data used only ~10 imbalanced classes. The author of the dataset provide:\r\n\r\n- a list of the 174 discourse markers\r\n- a Base version of the dataset with 1.74 million pairs (10k examples per marker)\r\n- a Big version with 3.4 million pairs\r\n- a Hard version with 1.74 million pairs where the connective couldn't be predicted with a fastText linear model\r\n\r\nSource: [GitHub](https://github.com/synapse-developpement/Discovery)","description_withheld":null,"homepage":"https://github.com/synapse-developpement/Discovery","introduced_date":"2019-03-28","introduced_date_note":null,"introduced_by":{"paper":"/paper/mining-discourse-markers-for-unsupervised","title":"Mining Discourse Markers for Unsupervised Sentence Representation Learning","first_author":"Damien Sileo","url":null},"license":{"name":"Apache 2.0","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Relation Classification","url":"/task/relation-classification","datasets_with_task":"/datasets/task/relation-classification"},{"name":"Sentence Embeddings","url":"/task/sentence-embeddings","datasets_with_task":"/datasets/task/sentence-embeddings"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Discovery"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/sileod/discovery","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/discovery","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/synapse-developpement/Discovery","url":"https://github.com/synapse-developpement/Discovery","frameworks":[]}],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/relation-classification-on-discovery-dataset","task":"Relation Classification","dataset_variant":"Discovery","rows":1,"metrics":["1:1 Accuracy"],"first_row_in_archive_order":{"model":"BERT","paper":"/paper/mining-discourse-markers-for-unsupervised","metrics":{"1:1 Accuracy":"20.6"},"code_links":[{"title":"synapse-developpement/Discovery","url":"https://github.com/synapse-developpement/Discovery"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/mining-discourse-markers-for-unsupervised","title":"Mining Discourse Markers for Unsupervised Sentence Representation Learning","date":"2019-03-28","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}