{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/help-a-dataset-for-identifying-shortcomings","title":"HELP: A Dataset for Identifying Shortcomings of Neural Models in Monotonicity Reasoning","arxiv_id":"1904.12166","date":"2019-04-27","proceeding":"SEMEVAL 2019 6","authors":["Hitomi Yanaka","Koji Mineshima","Daisuke Bekki","Kentaro Inui","Satoshi Sekine","Lasha Abzianidze","Johan Bos"],"abstract":"Large crowdsourced datasets are widely used for training and evaluating\nneural models on natural language inference (NLI). Despite these efforts,\nneural models have a hard time capturing logical inferences, including those\nlicensed by phrase replacements, so-called monotonicity reasoning. Since no\nlarge dataset has been developed for monotonicity reasoning, it is still\nunclear whether the main obstacle is the size of datasets or the model\narchitectures themselves. To investigate this issue, we introduce a new\ndataset, called HELP, for handling entailments with lexical and logical\nphenomena. We add it to training data for the state-of-the-art neural models\nand evaluate them on test sets for monotonicity phenomena. The results showed\nthat our data augmentation improved the overall accuracy. We also find that the\nimprovement is better on monotonicity inferences with lexical replacements than\non downward inferences with disjunction and modification. This suggests that\nsome types of inferences can be improved by our data augmentation while others\nare immune to it.","url_abs":"http://arxiv.org/abs/1904.12166v1","url_pdf":"http://arxiv.org/pdf/1904.12166v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"help-a-dataset-for-identifying-shortcomings","repo_url":"https://github.com/verypluming/HELP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"}],"methods":[],"datasets_introduced":[{"slug":"help","name":"HELP","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.12166","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1904.12166"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/verypluming/HELP","reach":null}],"summary":{"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"ca9955d1b0c834af","entry":"remove_duplicates","repo":"verypluming/HELP","repo_kind":"official","path":"scripts/create_dataset_PMB.py","file_url":"https://github.com/verypluming/HELP/blob/HEAD/scripts/create_dataset_PMB.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"CC-BY-SA-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"ca9955d1b0c834af"}},{"code_sha256_prefix":"04aceb0e02c82925","entry":"keep_tenses","repo":"verypluming/HELP","repo_kind":"official","path":"scripts/create_dataset_PMB.py","file_url":"https://github.com/verypluming/HELP/blob/HEAD/scripts/create_dataset_PMB.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"CC-BY-SA-4.0","inline_ok":false,"mcp_get_code":{"code_sha256":"04aceb0e02c82925"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}