Browse State-of-the-Art › Legal Reasoning
Legal Reasoning
30 papers with code · 2 benchmarks · 0 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| LegalBench (Issue-spotting) (3 rows) | GPT-4 | — | — | — | Compare |
| LegalBench (Rule-recall) (1 row) | GPT-4 | GPT-4 Technical Report | code | Syntology ran 2 of 5 samples · 3 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 30 papers with code (92 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
15 Mar 2023 11 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.
-
20 Sep 2023 2 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedWe propose DISC-LawLLM, an intelligent legal system utilizing large language models (LLMs) to provide a wide range of legal services.
-
15 Jun 2023 2 repositories listedOur benchmark contains diverse datasets from the Swiss legal system, allowing for a comprehensive study of the underlying non-English, inherently multilingual legal system.
-
2 Jun 2025 1 repository listedThis study uses Jordanian law as a case study to explore the fine-tuning of the Llama-3.
-
29 May 2025 1 repository listedHealth, Safety, and Environment (HSE) compliance assessment demands dynamic real-time decision-making under complicated regulations and complex human-machine-environment interactions.
-
4 May 2025 1 repository listedThis paper presents a domain-specific implementation of Retrieval-Augmented Generation (RAG) tailored to the Fair Use Doctrine in U.
-
25 Feb 2025 1 repository listedTo the best of our knowledge, CaseGen is the first benchmark designed to evaluate LLMs in the context of legal case document generation.
-
24 Feb 2025 1 repository listed Syntology ran 0 of 15 samples · 15 unverifiedThe Four-Element Theory is a fundamental framework in criminal law, defining the constitution of crime through four dimensions: Subject, Object, Subjective aspect, and Objective aspect.
-
15 Feb 2025 1 repository listedThe application of large language models (LLMs) in the legal domain holds significant potential for information retrieval and question answering, yet Thai legal QA systems face challenges due to a lack of standardized…
-
11 Feb 2025 1 repository listedLarge Language Models (LLMs) have achieved impressive results across numerous domains, yet they experience notable deficiencies in legal question-answering tasks.
-
10 Feb 2025 1 repository listedTo address these limitations, we study data generation for legal reasoning to improve the legal reasoning performance of open-source LLMs with the help of proprietary LLMs.
-
8 Feb 2025 1 repository listedWe then develop an LLM-based automated evaluation framework to identify reasoning errors and evaluate the performance of LLMs.
-
16 Oct 2024 1 repository listedWe remark that existing works investigate the phenomenon of weak-to-strong generation in analogous setup (i.
-
11 Oct 2024 1 repository listedHere, we introduce KBL, a benchmark for assessing the Korean legal language understanding of LLMs, consisting of (1) 7 legal knowledge tasks (510 examples), (2) 4 legal reasoning tasks (288 examples), and (3) the Korean…
-
3 Oct 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Large Language Models (LLMs) could struggle to fully understand legal theories and perform complex legal reasoning tasks.
-
19 Jul 2024 1 repository listedExisting benchmarks for evaluating knowledge update methods are mostly designed for the open domain and cannot address the specific challenges of the legal domain, such as the nuanced application of new legal knowledge,…
-
15 Apr 2024 1 repository listedThe recent proliferation of generative artificial intelligence (AI) technologies such as pre-trained large language models (LLMs) has opened up new frontiers in computational law.
-
2 Apr 2024 1 repository listedWe publish LawInstruct as a resource for further study of instruction tuning in the legal domain.
-
31 Mar 2024 1 repository listedIn common law jurisdictions, legal practitioners rely on precedents to construct arguments, in line with the doctrine of \emph{stare decisis}.
-
16 Feb 2024 1 repository listedWe introduce a new prompting method, Chain of Logic, which elicits rule-based reasoning through decomposition (solving elements as independent threads of logic), and recomposition (recombining these sub-answers to…
-
8 Dec 2023 1 repository listedAnnually, e-commerce platforms incur substantial financial losses due to trademark infringements, making it crucial to identify and mitigate potential legal risks tied to merchant information registered to the platforms.
-
27 Oct 2023 1 repository listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)Our findings generally sound a note of caution in the use of generative LMs on complex tasks without fine-tuning and point to the continued relevance of human annotation-intensive classification methods.
-
23 Oct 2023 1 repository listedEach scenario in the corpus is annotated with a complete IRAC analysis in a semi-structured format so that both machines and legal professionals are able to interpret and understand the annotations.
-
18 Oct 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Large language models (LLMs) have demonstrated great potential for domain-specific applications, such as the law domain.
-
LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models20 Aug 2023 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform?
-
13 Sep 2022 1 repository listedFinally-inspired by the Open Science movement-we make a call for the legal and computer science communities to join our efforts by contributing new tasks.
-
1 Jun 2022 1 repository listedWe investigate an automated approach to extract legal claims from news articles and to match the claims with their corresponding applicable laws.
-
25 Mar 2019 1 repository listedA framework and methodology---termed LogiKEy---for the design and engineering of ethical reasoners, normative theories and deontic logics is presented.
-
14 Dec 2017 1 repository listedIn Brazil, all legal professionals must demonstrate their knowledge of the law and its application by passing the OAB exams, the national bar exams.
-
29 Aug 2016 1 repository listedThe theory of actual causality, defined by Halpern and Pearl, and its quantitative measure - the degree of responsibility - was shown to be extremely useful in various areas of computer science due to a good match…
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections