Home › Code › process_dataset

process_dataset

Syntologyentry name in harvested coderead from the graph 2026-09-24

process_dataset appears in the code Syntology harvested for 20 papers, as 21 distinct code bodies found in 21 places (a place is one code body under one paper). At least one of them ran in 7 of the papers; 0 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named process_dataset do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 8 of the 21 distinct code bodies named process_dataset; 13 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
2ran · our draft was wrong
0ran · fixture could not drive it
6ran
13unverified
0fingerprinted

Licence is a property of each copy, so it is counted per place: 12 of the 21 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

20 papers shown of 20, newest first; 21 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 5 papers added by Syntology. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
BSC-Net: A Small-Branch-Sensitive Structural Continuity Network for Coronary Vessel Segmentation and Quantitative Angiographic Analysis added by Syntology 2026-09 (from id) liwx-deeplearning/BSC-Net/postprocessing/vessel_repair.py 55661f629f654c22 unverified MIT (permissive)
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction added by Syntology 2026-08 (from id) SetonLiang/Doc2DB-Bench/evaluation/clean_llamaextract_result.py ed2b9770a39a07ce unverified no licence file found · pointer only
It's Not What You Say, It's How You Say It: Evaluating LLM Responses to Expressions of Belief added by Syntology 2026-07 (from id) clarakuempel/EoB/score_dataset.py 9a42667a07ed158c unverified no licence file found · pointer only
Generating Concept Lexicalizations via Dictionary-Based Cross-Lingual Sense Projection added by Syntology 2026-04 (from id) UAlberta-NLP/ExpandNet/expandnet_step3_project.py db2b44164286fc61 ran · our draft was wrong no licence file found · pointer only
NO RELIABLE EVIDENCE OF SELF-REPORTED SENTIENCE IN LARGE LANGUAGE MODELS added by Syntology 2026-01 (from id) casparwarwick/sentient_machines_public/01_cluster_pipeline/get_continuation.py e7a7f254fe7d6701 unverified no licence file found · pointer only
Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents 19 May 2025 runamu/monday/data_processing/extract_scenes.py c9a868b25b751c41 unverified Apache-2.0 (permissive)
RARE: Retrieval-Augmented Reasoning Modeling 30 Mar 2025 open-dataflow/rare/process/process_medqa.py bf7b0dc382946a22 unverified Apache-2.0 (permissive)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers 23 Jan 2025 acharaakshit/FairMod/utils.py 832c2aa62967a652 unverified MIT (permissive)
Can Language Models Perform Robust Reasoning in Chain-of-thought Prompting with Noisy Rationales? 31 Oct 2024 dsubuntu/scie/prepare_data_2.py 1d31f1da783c7f25 ran · our draft was wrong no licence file found · pointer only
Embedding-based classifiers can detect prompt injection attacks 29 Oct 2024 AhsanAyub/malicious-prompt-detection/binary_classification.py b3a5bd4d1e163768 unverified no licence file found · pointer only
Beyond correlation: The Impact of Human Uncertainty in Measuring the Effectiveness of Automatic Evaluation and LLM-as-a-Judge 3 Oct 2024 amazon-science/beyondcorrelation/examples/judge_bench_example.py 992a1552125a9048 unverified Apache-2.0 (permissive)
UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models 27 Jun 2024 howiehwong/unigen/unigen/utils/challenge.py c81e8bf3d510cc9e unverified no licence file found · pointer only
Bayesian Online Natural Gradient (BONG) 30 May 2024 petergchang/bong/bong/src/dataloaders.py 03c0e3a373096d5d ran MIT (permissive)
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals 8 May 2024 sevdeawesome/POSER/src/detection_strategies/can_we_shift_it.py 7d71c3773a5cd44f ran no licence file found · pointer only
Towards Trustworthy Reranking: A Simple yet Effective Abstention Mechanism 20 Feb 2024 artefactory/abstention-reranker/abstention_reranker/utils.py f53c7ebd02857bf4 ran MIT (permissive)
LLM-as-a-Coauthor: Can Mixed Human-Written and Machine-Generated Text Be Detected? 11 Jan 2024 dongping-chen/mixset/dataset_loader.py d6120a808c697802 unverified no licence file found · pointer only
Learning to Abstract with Nonparametric Variational Information Bottleneck 26 Oct 2023 idiap/nvib_selfattention/data_modules/ArxivDataModule.py e2a2369f2cb401a7 ran no licence file found · pointer only
Learning to Abstract with Nonparametric Variational Information Bottleneck 26 Oct 2023 idiap/nvib_selfattention/data_modules/SentEvalDataModule.py 2c6d88e3f4e079ca ran no licence file found · pointer only
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing 2 Sep 2023 cwang621/blsp/blsp/src/text_instruction_dataset.py e35ba6d4c2745c08 ran Apache-2.0 (permissive)
CharacterChat: Learning towards Conversational AI with Personalized Social Support 20 Aug 2023 morecry/characterchat/model/BERT/util.py 6681fc262a4fb058 unverified no licence file found · pointer only
Listen, Attend and Spell 5 Aug 2015 WindQAQ/listen-attend-and-spell/utils/dataset_utils.py 00ec8a3f07cfb9b9 unverified Apache-2.0 (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections