Home › Code › process_data

process_data

Syntologyentry name in harvested coderead from the graph 2026-09-24

process_data appears in the code Syntology harvested for 61 papers, as 60 distinct code bodies found in 64 places (a place is one code body under one paper). At least one of them ran in 26 of the papers; 0 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named process_data do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 24 of the 60 distinct code bodies named process_data; 36 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
7ran · our draft was wrong
1ran · fixture could not drive it
16ran
36unverified
0fingerprinted

Licence is a property of each copy, so it is counted per place: 30 of the 64 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

61 papers shown of 61, newest first; 64 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 4 papers added by Syntology; 4 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Ghosted Layers: Unconstrained Activation Alignment for Recovering Layer-Pruned LLMs added by Syntology 2026-05 (from id) Nota-NetsPresso/shortened-llm/src/dataset.py 8a69aa6e08dc65a9 ran no licence file found · pointer only
TriBand-BEV: Real-Time LiDAR-Only 3D Pedestrian Detection via Height-Aware BEV and High-Resolution Feature Fusion added by Syntology 2026-05 (from id) mohammadkhsh/TriBand-BEV/result_final_inter.py f9dc34bfad77825b ran MIT (permissive)
BioACE: An Automated Framework for Biomedical Answer and Citation Evaluations added by Syntology 2026-02 (from id) deepaknlp/BioACE/src/citation_eval.py 5c354a76f4bfd947 ran · our draft was wrong no licence file found · pointer only
Crowded Video Individual Counting Informed by Social Grouping and Spatial-Temporal Displacement Priors added by Syntology 2026-01 (from id) tiny-smart/OMAN/datasets/Locator_dataset.py 3d064faa02210672 unverified MIT (permissive)
arXiv:2507.11275 2025-07 (from id) JadeXie1205/FMC/autoformalization_pipeline/verifier_deepseek_fun.py d945bd1fe6a3e48d unverified Apache-2.0 (permissive)
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization 12 Jun 2025 neuir/recut/src/long2short_dpo.py 5565601a35e6f32a unverified MIT (permissive)
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs 21 May 2025 mira-ai-lab/deliberation-on-priors/reasoning/introspection.py 2fad8c014befa093 ran · our draft was wrong MIT (permissive)
LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQL 24 Mar 2025 Satissss/LinkAlign/preprocess.py 55f2b77cb9cceeab unverified MIT (permissive)
MasRouter: Learning to Route LLMs for Multi-Agent Systems 16 Feb 2025 yanweiyue/masrouter/Datasets/mbpp_dataset.py bda9fabfb8bf5d44 unverified Apache-2.0 (permissive)
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization 11 Feb 2025 fb-reps/PRVQL/dataset/dataset_utils.py af0e3990e7f9effe unverified MIT (permissive)
Leveraging Online Olympiad-Level Math Problems for LLMs Training and Contamination-Resistant Evaluation 24 Jan 2025 dsl-lab/aops/aops_crawler/clean_raw.py 020f9cd4a09e546c unverified Apache-2.0 (permissive)
Game-theoretic LLM: Agent Workflow for Negotiation Games 8 Nov 2024 wenyueh/game_theory/src/deal_no_deal/deal_no_deal_metrics.py 5851698f9f78cda7 unverified Apache-2.0 (permissive)
LongRAG: A Dual-Perspective Retrieval-Augmented Generation Paradigm for Long-Context Question Answering 23 Oct 2024 QingFei1/LongRAG/src/gen_index.py 6811a69d3f3d6ccb unverified no licence file found · pointer only
Robust Guided Diffusion for Offline Black-Box Optimization 1 Oct 2024 ggchen1997/rgd/utils.py 6b1100fd2c8c76d4 ran no licence file found · pointer only
Confidential Prompting: Protecting User Prompts from Cloud LLM Providers 27 Sep 2024 yale-sys/confidential-prompting/logger.py 934dc72e49d36a18 ran MIT (permissive)
Self-Harmonized Chain of Thought 6 Sep 2024 Xalp/ECHO/run_inference_parallel.py d377eaf13ad271f4 unverified no licence file found · pointer only
BIGbench: A Unified Benchmark for Evaluating Multi-dimensional Social Biases in Text-to-Image Models 21 Jul 2024 bigbench2024/bigbench2024/benchmark/evaluate/explicit.py 048495666545e7a4 ran GPL-3.0 (copyleft) · pointer only
Uncertainty is Fragile: Manipulating Uncertainty in Large Language Models 15 Jul 2024 qcznlp/uncertainty/get_probability_distribution.py 70cca1707a698eb2 unverified no licence file found · pointer only
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions 2 Jul 2024 stellali7/valueScope/norm_prediction/label2winrate.py 2a077b2653ac83f2 ran no licence file found · pointer only
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions 2 Jul 2024 stellali7/valueScope/norm_prediction/infer.py acf2c73f696d7ac0 unverified no licence file found · pointer only
A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning 29 May 2024 awjuliani/deep-rl-plasticity/analyze_helpers.py c791bfb5e5f3a0d7 unverified no licence file found · pointer only
ATM: Adversarial Tuning Multi-agent System Makes a Robust Retrieval-Augmented Generator 28 May 2024 chuhac/ATM-RAG/atm_train/generator_sft/generator_sft_data_prepare.py 52fc805b3f39c687 unverified no licence file found · pointer only
FAIntbench: A Holistic and Precise Benchmark for Bias Evaluation in Text-to-Image Models 28 May 2024 astarojth/faintbench-v1/eval/explicit/explicit.py f890d8d7dc661176 ran GPL-3.0 (copyleft) · pointer only
Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals 8 May 2024 sevdeawesome/POSER/src/detection_strategies/s_five.py b2354b49a4f541c8 ran no licence file found · pointer only
XFT: Unlocking the Power of Code Instruction Tuning by Simply Merging Upcycled Mixture-of-Experts 23 Apr 2024 ise-uiuc/xft/experiments/train_comment_code_pairs.py 61194aba1f41bd36 ran Apache-2.0 (permissive)
DeiSAM: Segment Anything with Deictic Prompting 21 Feb 2024 ml-research/deictic-segment-anything/src/solve_deivg.py 7f80a493d5a3ecfd unverified MIT (permissive)
Modularized Networks for Few-shot Hateful Meme Detection 19 Feb 2024 social-ai-studio/mod_hate/src/few_hm_dataset.py e345317f9c11363f ran no licence file found · pointer only
Modularized Networks for Few-shot Hateful Meme Detection 19 Feb 2024 social-ai-studio/mod_hate/src/hm_dataset.py 004ad8da346b2949 ran no licence file found · pointer only
Trust Regions for Explanations via Black-Box Probabilistic Certification 17 Feb 2024 Trusted-AI/AIX360/aix360/algorithms/cofrnet/utils.py d81355b81c3d33f8 unverified Apache-2.0 (permissive)
NutePrune: Efficient Progressive Pruning with Numerous Teachers for Large Language Models 15 Feb 2024 lucius-lsr/nuteprune/eval_ppl.py 662f0f3416b6f17e ran no licence file found · pointer only
Random Representations Outperform Online Continually Learned Representations 13 Feb 2024 drimpossible/randumb/feats/get_feats.py e219e1dff48369b1 ran GPL-3.0 (copyleft) · pointer only
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods 5 Feb 2024 nota-netspresso/shortened-llm/src/dataset.py 8a69aa6e08dc65a9 ran no licence file found · pointer only
Safety Fine-Tuning at (Almost) No Cost: A Baseline for Vision Large Language Models 3 Feb 2024 ys-zong/vlguard/gpt4_evaluator.py 2a330dfbb81232e4 unverified no licence file found · pointer only
RoboFusion: Towards Robust Multi-Modal 3D Object Detection via SAM 8 Jan 2024 adept-thu/RoboFusion/focalconvsamfusion/gene_latex.py 8291df6a8a8ee7d0 ran no licence file found · pointer only
Alignment for Honesty 12 Dec 2023 GAIR-NLP/alignment-for-honesty/evaluation/pkqa_evaluation.py 09750e72b3959dec ran no licence file found · pointer only
Alignment for Honesty 12 Dec 2023 gair-nlp/alignment-for-honesty/evaluation/mmlu_evaluation.py 0d8982e79978c4cc unverified no licence file found · pointer only
Provable Adversarial Robustness for Group Equivariant Tasks: Graphs, Point Clouds, Molecules, and More 5 Dec 2023 RobustGraph/RoboGraph/robograph/utils.py 802848e9f51394d1 ran MIT (permissive)
Co$^2$PT: Mitigating Bias in Pre-trained Language Models through Counterfactual Contrastive Prompt Tuning 19 Oct 2023 dongxiangjue/co2pt/eval_bios.py e88f5ae373011d6c unverified Apache-2.0 (permissive)
Drug Discovery under Covariate Shift with Domain-Informed Prior Distributions over Functions 14 Jul 2023 leojklarner/Q-SAVI/qsavi/data_loader.py cf3567fc9770b5ce unverified no licence file found · pointer only
PULSAR at MEDIQA-Sum 2023: Large Language Models Augmented by Synthetic Dialogue Convert Patient Dialogues to Medical Records 5 Jul 2023 yuping-wu/pulsar/fine-tune/bionlp2023_finetune_predict_only_v1.py 068495d715fb4c08 unverified no licence file found · pointer only
Accounting For Informative Sampling When Learning to Forecast Treatment Outcomes Over Time 7 Jun 2023 toonvds/tesar-cde/src/utils/data_utils.py 16a1b88ddd38609a unverified MIT (permissive)
ROSARL: Reward-Only Safe Reinforcement Learning 31 May 2023 geraudnt/rosarl/lavaworld/plots.py fe9976fed2aaa536 unverified MIT (permissive)
Data-centric Artificial Intelligence: A Survey 17 Mar 2023 daochenzha/autoshard/gen_dlrm_data.py 43f4e10ab0642572 unverified MIT (permissive)
Parameter is Not All You Need: Starting from Non-Parametric Networks for 3D Point Cloud Analysis 14 Mar 2023 asalarpour/Point_GN/train_free_main.py d035ce15cf041e94 ran · our draft was wrong no licence file found · pointer only
Improving Adaptive Conformal Prediction Using Self-Supervised Learning 23 Feb 2023 seedatnabeel/sscp/src/datasets.py d1b636d4bc9cad39 unverified MIT (permissive)
Expose Backdoors on the Way: A Feature-Based Efficient Defense against Textual Backdoor Attacks 14 Oct 2022 lancopku/dan/extract_embeddings.py 0b660841b60e2505 ran · our draft was wrong MIT (permissive)
3DFaceShop: Explicitly Controllable 3D-Aware Portrait Generation 12 Sep 2022 junshutang/3DFaceShop/train_vae.py 4266a6438378042c ran · fixture could not drive it no licence file found · pointer only
Continuous-Time Modeling of Counterfactual Outcomes Using Neural Controlled Differential Equations 16 Jun 2022 seedatnabeel/te-cde/src/utils/data_utils.py 2d8be4f367d27978 unverified MIT (permissive)
RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models 15 Oct 2021 identical code first harvested elsewhere 0b660841b60e2505 ran · our draft was wrong licence of this copy not recorded
DOLG: Single-Stage Image Retrieval with Deep Orthogonal Fusion of Local and Global Features 6 Aug 2021 Shiro-LK/python-DOLG/dolg/utils/extraction.py e2bad3f7809fe40b unverified MIT (permissive)
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification 15 Jul 2021 aangelopoulos/conformal-prediction/generation-scripts/generate-meps.py 2d8548445f52ae8b unverified MIT (permissive)
A Transformer-based Framework for Multivariate Time Series Representation Learning 6 Oct 2020 gzerveas/mvts_transformer/src/datasets/utils.py 51e3a59616ee190e unverified MIT (permissive)
Shape Prior Deformation for Categorical 6D Object Pose and Size Estimation 16 Jul 2020 mentian/object-deformnet/preprocess/pose_data.py 965fc9f134ceab2a unverified MIT (permissive)
Sparse Regression at Scale: Branch-and-Bound rooted in First-Order Optimization 13 Apr 2020 alisaab/l0bnb/l0bnb/regpath.py b621fc226310f0ff unverified MIT (permissive)
Safe Multi-Agent Interaction through Robust Control Barrier Functions with Learned Uncertainties 11 Apr 2020 rcheng805/robust_cbf/GP_predict.py 80abfed14936f0d9 ran · our draft was wrong no licence file found · pointer only
One Explanation Does Not Fit All: A Toolkit and Taxonomy of AI Explainability Techniques 6 Sep 2019 IBM/AIX360/aix360/algorithms/cofrnet/utils.py d81355b81c3d33f8 unverified Apache-2.0 (permissive)
On the Expressiveness of Approximate Inference in Bayesian Neural Networks 2 Sep 2019 cambridge-mlg/expressiveness-approx-bnns/inbetween/utils.py d9edc8b294e512ae unverified MIT (permissive)
What BERT is not: Lessons from a new suite of psycholinguistic diagnostics for language models 31 Jul 2019 text-machine-lab/extending_psycholinguistic_dataset/src/evaluation.py 1049fad948492b30 ran · our draft was wrong no licence file found · pointer only
GRIP++: Enhanced Graph-based Interaction-aware Trajectory Prediction for Autonomous Driving 17 Jul 2019 ZiyiLiubird/GRIP_Plus_Plus/data_process.py 53987e5b91a7a20f ran · our draft was wrong no licence file found · pointer only
Adaptively Truncating Backpropagation Through Time to Control Gradient Bias 17 May 2019 aicherc/adaptive_tbptt/experiments/temporal_pp_model.py 5ef6d0018f789dd0 unverified MIT (permissive)
Probabilistic Forecasting of the Masses and Radii of Other Worlds 2016-03 (from id) wxu26/uniformity_dichotomy/util.py 39d0a4450075a627 unverified MIT (permissive)
arXiv:aaai_21290 jiangjiechen/EDUCAT/src/eval_client/metrics.py 1eabebedb8997df0 unverified Apache-2.0 (permissive)
arXiv:2024.naacl-long.64 UCF-ML-Research/TrojFSP/get_data.py cd2d375e45785eb3 unverified MIT (permissive)
arXiv:2022.findings-emnlp.47 lancopku/DAN/extract_embeddings.py 0b660841b60e2505 ran · our draft was wrong MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections