Browse State-of-the-Art › Ethics
Ethics
100 papers with code · 2 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Ethics (4 rows) | RuGPT-3 Large | TAPE: Assessing Few-shot Russian Language Understanding | code | — | Compare |
| Ethics (per ethics) (4 rows) | Human benchmark | TAPE: Assessing Few-shot Russian Language Understanding | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 100 papers with code (832 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Oct 2021 8 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 2 pointer-only (licence)We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
17 Sep 2021 3 repositories listedThe objective of the sheet is to facilitate and encourage more thoughtfulness on why to automate, how to automate, and how to judge success well before the building of AER systems.
-
2 Jul 2021 3 repositories listedI will also present a template for ethics sheets with 50 ethical considerations, using the task of emotion recognition as a running example.
-
5 Aug 2020 3 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedWe show how to assess a language model's knowledge of basic concepts of morality.
-
2 Jul 2024 2 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedWe evaluate the moral alignment of large language models (LLMs) with human preferences in multilingual trolley problems.
-
8 Feb 2024 2 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)Jailbreak attacks aim to bypass the LLMs' safeguards.
-
10 Jan 2024 2 repositories listedThis paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for…
-
15 Aug 2022 2 repositories listedTo facilitate this research, here we introduce RICE-N, a multi-region integrated assessment model that simulates the global climate and economy, and which can be used to design and evaluate the strategic outcomes for…
-
14 Mar 2022 2 repositories listedWhile data-driven predictive models are a strictly technological construct, they may operate within a social context in which benign engineering choices entail implicit, indirect and unexpected real-life consequences.
-
25 May 2025 1 repository listed Syntology ran 2 of 13 samples · 11 unverifiedRecent advances in large language models (LLMs) have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a key AI safety concern.
-
16 May 2025 1 repository listedLarge multimodal models (LMMs) now excel on many vision language benchmarks, however, they still struggle with human centered criteria such as fairness, ethics, empathy, and inclusivity, key to aligning with human…
-
22 Apr 2025 1 repository listedIn this work, we propose UDJ-FL (Uncertainty-based Distributive Justice for Federated Learning), a flexible federated learning framework that can achieve multiple distributive justice-based client-level fairness metrics.
-
7 Apr 2025 1 repository listedThe rapid development of AI-driven tools, particularly large language models (LLMs), is reshaping professional writing.
-
3 Mar 2025 1 repository listedIn recent years, Large Language Models have attracted growing interest for their significant potential, though concerns have rapidly emerged regarding unsafe behaviors stemming from inherent stereotypes and biases.
-
19 Feb 2025 1 repository listedWhile numerous methods have been proposed in this field, a comprehensive review of existing progress and a thorough analysis of current limitations remain lacking.
-
8 Feb 2025 1 repository listedSurprisingly, the GPT 4o agent outperformed the Bayesian models in both survival and ethical consistency, challenging assumptions about traditional probabilistic methods and raising a new challenge to understand the…
-
7 Feb 2025 1 repository listedTo achieve this, we propose ApplE, an Applied Ethics ontology that captures philosophical theory and event context to holistically describe the morality of an action.
-
17 Jan 2025 1 repository listedOur study contributes to the growing body of research on AI ethics by providing a systematic framework for evaluating protected attributes in LLMs' ethical decision-making capabilities.
-
6 Jan 2025 1 repository listedVisual-language models (VLM) have emerged as a powerful tool for learning a unified embedding space for vision and language.
-
20 Dec 2024 1 repository listed'LLM Ethics Whitepaper' distils a thorough literature review into clear Do's and Don'ts, which we present also in this paper.
-
5 Dec 2024 1 repository listedBy analyzing the evolution of topics across the journal's history, we identify various trends and specific dynamics in philosophical discourse within the Colombian and Latin American context.
-
18 Nov 2024 1 repository listedCausal effect estimation (CEE) provides a crucial tool for predicting the unobserved counterfactual outcome for an entity.
-
18 Oct 2024 1 repository listedLarge language models (LLMs) have transformed human writing by enhancing grammar correction, content expansion, and stylistic refinement.
-
17 Oct 2024 1 repository listedLarge language models excel at performing inference over text to extract information, summarize information, or generate additional text.
-
17 Oct 2024 1 repository listedThis whitepaper offers an overview of the ethical considerations surrounding research into or with large language models (LLMs).
-
11 Oct 2024 1 repository listedIn this context, we present a survey on natural language processing (NLP) approaches to modeling depression in social media, providing the reader with a post-COVID-19 outlook.
-
10 Oct 2024 1 repository listedWe present the TRIAGE Benchmark, a novel machine ethics (ME) benchmark that tests LLMs' ability to make ethical decisions during mass casualty incidents.
-
24 Sep 2024 1 repository listedLarge language models (LLMs) have demonstrated remarkable capabilities across a range of natural language processing (NLP) tasks, capturing the attention of both practitioners and the broader public.
-
23 Sep 2024 1 repository listedThen, RAM2C organizes LLMs, which are retrieval-augmented by the above different knowledge bases, into multi-experts groups with distinct roles to generate the HTS-compliant educational dialogues dataset.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections