Home › Code › prepare_dataset

prepare_dataset

Syntologyentry name in harvested coderead from the graph 2026-09-24

prepare_dataset appears in the code Syntology harvested for 32 papers, as 30 distinct code bodies found in 33 places (a place is one code body under one paper). At least one of them ran in 5 of the papers; 0 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named prepare_dataset do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 6 of the 30 distinct code bodies named prepare_dataset; 24 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
0ran · our draft was wrong
0ran · fixture could not drive it
6ran
24unverified
0fingerprinted

Licence is a property of each copy, so it is counted per place: 11 of the 33 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

32 papers shown of 32, newest first; 33 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 5 papers added by Syntology; 2 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Cognition on Graph: Navigating Massive Knowledge Space via Cognitive Cycles and Bidirectional Graph-Text Synergy added by Syntology 2026-09 (from id) zhougengxian/CoG/CoG/utils.py 7d304feda300ff3f unverified Apache-2.0 (permissive)
Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines added by Syntology 2026-05 (from id) StellarLuminosity/Energy/distill_bench/core/utils.py 7aca9ebb37ee1d3e ran no licence file found · pointer only
A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation added by Syntology 2026-05 (from id) sjmeis/DP-Decompose-Distribute/eval_scripts/adaptive_attacker.py 87cd9dd4f015d968 ran MIT (permissive)
A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation added by Syntology 2026-05 (from id) sjmeis/DP-Decompose-Distribute/eval_scripts/baseline_attacker.py 2927da7397153f77 ran MIT (permissive)
2 RELATED WORK Reinforcement learning has emerged as a powerful paradigm for enhancing the reasoning abilities of LLMs added by Syntology 2026-01 (from id) zywang0104/Video-KTR/src/r1-v/src/open_r1/sft_video.py 69e027783b4d6f9f unverified no licence file found · pointer only
SPoRC-VIST: A Benchmark for Evaluating Generative Natural Narrative in Vision-Language Models added by Syntology 2026-01 (from id) Yunlin-Zeng/visual-podcast-VLM/scripts_235B/finetune_235B.py a5c5ad9dfba730d2 unverified licence not identified · pointer only
arXiv:2507.01299 2025-07 (from id) alibaba/EfficientAI/masquant/custom_dataset.py d8f401d3be5d49f4 unverified Apache-2.0 (permissive)
EDINET-Bench: Evaluating LLMs on Complex Financial Tasks using Japanese Financial Statements 10 Jun 2025 sakanaai/edinet-bench/src/edinet_bench/logistic.py 67b49b0983888fd3 unverified Apache-2.0 (permissive)
VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning 18 May 2025 qiwang98/videorft/src/r1-v/src/open_r1/sft_video.py 40640dbb4bed6b2c unverified Apache-2.0 (permissive)
ESC: Erasing Space Concept for Knowledge Deletion 3 Apr 2025 KU-VGI/ESC/unlearn.py e98c82942ca364b2 unverified no licence file found · pointer only
Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training 24 Mar 2025 bbartoldson/TBA/tba_tldr.py a2e246bb45f45630 unverified Apache-2.0 (permissive)
RILQ: Rank-Insensitive LoRA-based Quantization Error Compensation for Boosting 2-bit Large Language Model Accuracy 2 Dec 2024 aiha-lab/rilq/rilq_utils/data.py 2379dffb2ba023d0 unverified Apache-2.0 (permissive)
Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models 23 Oct 2024 mnoukhov/async_rlhf/online_dpo.py a2e246bb45f45630 unverified Apache-2.0 (permissive)
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 18 Oct 2024 SteveTANTAN/PoG/PoG/utils.py 68fca283c8ee65c1 unverified Apache-2.0 (permissive)
Neural Solver Selection for Combinatorial Optimization 13 Oct 2024 lamda-bbo/neural-solver-selection/utils.py 2eeabea8e2f0857a unverified MIT (permissive)
LLM-DetectAIve: a Tool for Fine-Grained Machine-Generated Text Detection 8 Aug 2024 mbzuai-nlp/llm-detectaive/pipeline/dataset.py d7175d9e4888d2f7 ran no licence file found · pointer only
Generate-on-Graph: Treat LLM as both Agent and KG in Incomplete Knowledge Graph Question Answering 23 Apr 2024 YaooXu/GoG/CoT/utils.py 00b2e6d240c0d36b unverified no licence file found · pointer only
Mitigating Privacy Risk in Membership Inference by Convex-Concave Loss 8 Feb 2024 ml-stat-Sustech/ConvexConcaveLoss/source/data_preprocessing/dataset_preprocessing.py 0f643892107414ea ran no licence file found · pointer only
Through the Dual-Prism: A Spectral Perspective on Graph Data Augmentation for Graph Classification 18 Jan 2024 yu-rp/dualprism/src/utils.py cbf4d0892481c707 ran no licence file found · pointer only
Chat Vector: A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New Languages 7 Oct 2023 augmxnt/deccp/abliterator.py d53507b37cb81222 unverified Apache-2.0 (permissive)
Deciphering Raw Data in Neuro-Symbolic Learning with Provable Guarantees 21 Aug 2023 abductivelearning/abl-tl/data.py 3cd8f7c384d495a9 unverified no licence file found · pointer only
Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph 15 Jul 2023 IDEA-FinAI/ToG/CoT/utils.py 00b2e6d240c0d36b unverified no licence file found · pointer only
Direct Preference Optimization: Your Language Model is Secretly a Reward Model 29 May 2023 kidist-amde/ddro-direct-document-relevance-optimization/src/pretrain/finetune_docTTTTTquery_NQ.py 19b1210f13924026 unverified Apache-2.0 (permissive)
A Gold Standard Dataset for the Reviewer Assignment Problem 23 Mar 2023 niharshah/goldstandard-reviewer-paper-match/scripts/prepare_dataset.py 05393c48abdec847 unverified no licence file found · pointer only
AdvMIL: Adversarial Multiple Instance Learning for the Survival Analysis on Whole-Slide Images 13 Dec 2022 liupei101/advmil/dataset/utils.py 6c12bc9946ecffcb unverified MIT (permissive)
Unsupervised Learning of Temporal Abstractions with Slot-based Transformers 25 Mar 2022 agopal42/slottar/preprocess.py ca4ef00b5db7d896 unverified MIT (permissive)
LAFITE: Towards Language-Free Training for Text-to-Image Generation 27 Nov 2021 oxygenlu/ratlip/code/lib/perpare.py d4c40e264c45657c unverified MIT (permissive)
Learning Implicit Sentiment in Aspect-based Sentiment Analysis with Supervised Contrastive Pre-Training 3 Nov 2021 Tribleave/SCAPT-ABSA/train/misc.py b9a0638f8e3d830b unverified MIT (permissive)
DAG Card is the new Model Card 24 Oct 2021 jacopotagliabue/metaflow-intent-prediction/local_flow/intent/src/prepare_dataset.py ec5dcfa1a6b0f7af unverified MIT (permissive)
You Do Not Need a Bigger Boat: Recommendations at Reasonable Scale in a (Mostly) Serverless and Open Stack 15 Jul 2021 jacopotagliabue/you-dont-need-a-bigger-boat/local_flow/intent/src/prepare_dataset.py ec5dcfa1a6b0f7af unverified MIT (permissive)
Conditional Channel Gated Networks for Task-Aware Continual Learning 31 Mar 2020 lit-leo/cgate/src/data.py 9dce6d4bf3b88343 unverified MIT (permissive)
Fine-tuning CNN Image Retrieval with No Human Annotation 3 Nov 2017 nikosefth/freedom/utils_retrieval.py 4da573acef3a31a2 unverified MIT (permissive)
arXiv:2025.findings-emnlp.538 SOLAR2025ARR/SOLAR/datagen/llm_rerank.py 0588343b65e27435 unverified MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections