Browse State-of-the-Art › De-identification
De-identification
54 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
De-identification is the task of detecting privacy-related entities in text, such as person names, emails and contact data.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 54 papers with code (174 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
13 Oct 2021 8 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 2 pointer-only (licence)We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite.
-
15 May 2020 3 repositories listedLearning disentangled representations of data is a fundamental problem in artificial intelligence.
-
19 Sep 2019 3 repositories listedFinally, we discuss the privacy concerns associated with sharing synthetic data produced by GANs and test their ability to withstand a simple membership inference attack.
-
6 Apr 2019 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Contextual word embedding models such as ELMo (Peters et al., 2018) and BERT (Devlin et al., 2018) have dramatically improved performance for many natural language processing (NLP) tasks in recent months.
-
16 Oct 2024 2 repositories listedMedical data employed in research frequently comprises sensitive patient health information (PHI), which is subject to rigorous legal frameworks such as the General Data Protection Regulation (GDPR) or the Health…
-
2 Jun 2024 2 repositories listedTo make medical datasets accessible without sharing sensitive patient information, we introduce a novel end-to-end approach for generative de-identification of dynamic medical imaging data.
-
25 Jan 2022 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedWe present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods.
-
30 Aug 2020 2 repositories listedThe proliferation of speech technologies and rising privacy legislation calls for the development of privacy preservation solutions for speech applications.
-
15 Jun 2025 1 repository listedMany models are pretrained on redacted text for privacy reasons.
-
13 Feb 2025 1 repository listedSynthetic datasets were compared to real-world neurosurgical data to assess fidelity (means, proportions, distributions, and bivariate correlations), utility (ML classifier performance on RWD), and privacy (duplication…
-
11 Feb 2025 1 repository listedThis study presents the first systematic attempt to identify prompts and text tokens in MIMIC-CXR that contribute the most to training data memorization.
-
16 Jan 2025 1 repository listedIn recent years, there has been an increasing interest in image anonymization, particularly focusing on the de-identification of faces and individuals.
-
14 Jan 2025 1 repository listedOur evaluation on two public datasets, CRAPII and TSCC, demonstrates that the fine-tuned GPT-4o-mini model achieves superior performance, with a recall of 0.
-
14 Dec 2024 1 repository listedFace de-identification (DeID) has been widely studied for common scenes, but remains under-researched for medical scenes, mostly due to the lack of large-scale patient face datasets.
-
10 Nov 2024 1 repository listedWe believe this work provides a path forward for (i) the release of large-scale synthetic patient message datasets that are stylistically similar to ground-truth samples and (ii) HIPAA-friendly data generation which…
-
24 Oct 2024 1 repository listedWe evaluated the approaches on publicly available privacy and VAD data sets to examine the strengths and weaknesses of the different anonymization techniques and highlight the promising efficacy of our approach.
-
2 Oct 2024 1 repository listedDe-identification is important in protecting patients' privacy for healthcare text analytics.
-
15 Sep 2024 1 repository listedThe system preserves critical medical information while introducing diversity in the generations and minimising re-identification risk.
-
8 Jul 2024 1 repository listedThe consequences of a healthcare data breach can be devastating for the patients, providers, and payers.
-
24 Jun 2024 1 repository listedMany privacy-preserving techniques have been proposed for this task aiming to transform the data while ensuring the privacy of individuals.
-
29 May 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)To address this, CheXpert Plus serves as a new collection of radiology data sources, made publicly available to enhance the scaling, performance, robustness, and fairness of models for all subsequent machine learning…
-
29 Apr 2024 1 repository listed[Abridged Abstract] Recent technological advances underscore labor market dynamics, yielding significant consequences for employment prospects and increasing job vacancy data across platforms and languages.
-
15 Mar 2024 1 repository listedFace de-identification in videos is a challenging task in the domain of computer vision, primarily used in privacy-preserving applications.
-
28 Feb 2024 1 repository listedOn the DAIC-WOZ dataset with ComparE16 features and an LSTM-only model, our method achieves an F1-Score of 0.
-
23 Oct 2023 1 repository listed Syntology ran 7 of 11 samples · 4 unverified · 11 pointer-only (licence)The proposed diffusion-model-based method can reliably and efficiently generate synthetic EHR time series, which facilitates the downstream medical data analysis.
-
26 May 2023 1 repository listedTo evaluate the performance of the proposed de-identification tool, a comparative study was conducted between several existing defacing and refacing tools, with two different segmentation algorithms (FAST and Morphobox).
-
20 May 2023 1 repository listedThe increasing number of benchmarks for Natural Language Processing (NLP) tasks in the computational job market domain highlights the demand for methods that can handle job-related tasks such as skill extraction, skill…
-
18 May 2023 1 repository listedData sharing is crucial for open science and reproducible research, but the legal sharing of clinical data requires the removal of protected health information from electronic health records.
-
20 Mar 2023 1 repository listedWhile prior works have explored image de-identification strategies based on synthetic averaging of images in other domains (e.
-
20 Mar 2023 1 repository listedThe digitization of healthcare has facilitated the sharing and re-using of medical data but has also raised concerns about confidentiality and privacy.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections