{"url":"/task/valnov","name":"ValNov","slug":"valnov","description_markdown":"Given a textual premise and conclusion candidate, the Argument-Validity-and-Novelty-Prediction-Shared-Task ValNov consists in predicting two aspects of a conclusion: its validity and novelty.\r\n\r\nValidity is defined as the degree to which the conclusion is justified with respect to the given premise. A conclusion is considered to be valid if it is supported by inferences that link the premise to the conclusion, based on logical principles or commonsense or world knowledge, which may be defeasible. A conclusion will be trivially considered valid if it repeats or summarizes the premise – in which case it can hardly be considered as novel.\r\n\r\nNovelty defines the degree to which the conclusion contains content that is new in relation to the premise. As extreme cases, a conclusion candidate that repeats or summarizes the premise or is unrelated to the premise will not be considered novel.","categories":[{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":2,"papers_with_code":1,"benchmarks":2,"benchmark_tables_in_archive":2,"benchmark_tables_shown":2,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":2,"subtasks":0,"parent_tasks":1},"benchmarks":[{"leaderboard":"/sota/valnov-on-valnov-subtask-a","slug":"valnov-on-valnov-subtask-a","dataset":"ValNov Subtask A","dataset_url":"/dataset/valnov-subtask-a","rows_in_archive":8,"metrics":["JOINT-F1","VAL-F1","NOV-F1"],"first_row_in_archive_order":{"model":"CLTeamL-3","paper_title":"Overview of the 2022 Validity and Novelty Prediction Shared Task","paper_url":"/paper/overview-of-the-2022-validity-and-novelty","paper_date":"","arxiv_id":null,"code_links":[{"title":"phhei/argsvalidnovel","url":"https://github.com/phhei/argsvalidnovel"}],"syntology":null}},{"leaderboard":"/sota/valnov-on-valnov-subtask-b","slug":"valnov-on-valnov-subtask-b","dataset":"ValNov Subtask B","dataset_url":"/dataset/valnov-subtask-b","rows_in_archive":3,"metrics":["JOINT-F1","VAL-F1","NOV-F1"],"first_row_in_archive_order":{"model":"NLP@UIT","paper_title":"Overview of the 2022 Validity and Novelty Prediction Shared Task","paper_url":"/paper/overview-of-the-2022-validity-and-novelty","paper_date":"","arxiv_id":null,"code_links":[{"title":"phhei/argsvalidnovel","url":"https://github.com/phhei/argsvalidnovel"}],"syntology":null}}],"datasets":[{"url":"/dataset/valnov-subtask-a","name":"ValNov Subtask A","full_name":"","num_papers_in_archive":2},{"url":"/dataset/valnov-subtask-b","name":"ValNov Subtask B","full_name":"","num_papers_in_archive":1}],"subtasks":[],"parent_tasks":[{"url":"/task/argument-mining","name":"Argument Mining"}],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":1,"of":1,"tagged_in_all":2,"items":[{"url":"/paper/overview-of-the-2022-validity-and-novelty","title":"Overview of the 2022 Validity and Novelty Prediction Shared Task","date":"2022-10-01","arxiv_id":null,"repositories_listed":1,"syntology":null}],"syntology_records":0,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}