{"url":"/dataset/masc","name":"MASC","full_name":"Manually Annotated Sub-Corpus","description_markdown":"The Manually Annotated Sub-Corpus (MASC) consists of approximately 500,000 words of contemporary American English written and spoken data drawn from the Open American National Corpus (OANC).\r\n\r\nAll of MASC includes manually validated annotations for sentence boundaries, token, lemma and POS;\r\nnoun and verb chunks; named entities (person, location, organization, date); Penn Treebank syntax;\r\ncoreference; and discourse structure.\r\n\r\nAdditional manually produced or validated annotations have been produced by the MASC project\r\nfor portions of the sub-corpus, including full-text annotation for FrameNet frame elements\r\nand a 100K+ sentence corpus with WordNet 3.1 sense tags, of which one-tenth are also annotated for\r\nFrameNet frame elements.\r\n\r\nAnnotations of all or portions of the sub-corpus for a wide variety of other linguistic phenomena\r\nhave been contributed by other projects, including PropBank, TimeBank, Pittsburgh opinion, and several others.\r\n\r\nUnlike most freely available corpora including a wide variety of linguistic annotations,\r\nMASC contains a balanced selection of texts from a broad range of genres.","description_withheld":null,"homepage":"https://www.anc.org/data/masc/","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":{"name":"Creative Commons Attribution 3.0 United States License.","url":null},"modalities":[],"tasks":[{"name":"Named Entity Recognition (NER)","url":"/task/named-entity-recognition-ner","datasets_with_task":"/datasets/task/named-entity-recognition-ner"},{"name":"Automatic Speech Recognition","url":"/task/automatic-speech-recognition-2","datasets_with_task":"/datasets/task/automatic-speech-recognition-2"},{"name":"Coreference Resolution","url":"/task/coreference-resolution","datasets_with_task":"/datasets/task/coreference-resolution"},{"name":"Part-Of-Speech Tagging","url":"/task/part-of-speech-tagging","datasets_with_task":"/datasets/task/part-of-speech-tagging"},{"name":"Constituency Parsing","url":"/task/constituency-parsing","datasets_with_task":"/datasets/task/constituency-parsing"},{"name":"Sentence segmentation","url":"/task/sentence-segmentation","datasets_with_task":"/datasets/task/sentence-segmentation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["MASC"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}