Home › Code › tokenize

tokenize

Syntologyentry name in harvested coderead from the graph 2026-09-24

tokenize appears in the code Syntology harvested for 176 papers, as 147 distinct code bodies found in 184 places (a place is one code body under one paper). At least one of them ran in 58 of the papers; 14 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named tokenize do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 48 of the 147 distinct code bodies named tokenize; 99 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

0ran · honoured contract
0ran · violated contract
17ran · our draft was wrong
0ran · fixture could not drive it
31ran
99unverified
14fingerprinted

Licence is a property of each copy, so it is counted per place: 62 of the 184 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

100 papers shown of 176 (newest first; the JSON twin carries all 184 places); 103 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 33 papers added by Syntology; 14 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Option-Aware Retrieval and Task-Specific VLM Adaptation for Medical VQA added by Syntology 2026-09 (from id) Kirscher/MedReason2026/medreason_baseline/retrieval.py 39f7fdbe9ffb0f20 unverified Apache-2.0 (permissive)
Solar Intelligence added by Syntology 2026-09 (from id) jyotsnasingh11217/Solar-Intelligence/build_bm25.py 8d254f76a316ef23 unverified licence not identified · pointer only
In the Blind: Building Pseudo-References for MT Evaluation added by Syntology 2026-09 (from id) surrey-nlp/PseudoRef/pseudoref/diagnostics.py 7a46793afd47ec22 unverified licence not identified · pointer only
Occlusal Geometry in Closed Form for Orthodontic Report Generation added by Syntology 2026-09 (from id) GIND123/ODIN_toothfairy4/src/bite2text/eval/gc_metrics.py bb2d7f0fe0f03ad9 unverified no licence file found · pointer only
EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision added by Syntology 2026-09 (from id) 18277390221/EmoStance/src/latent_stance_control/evaluate_text_quality.py c6e3fb7b29530122 unverified MIT (permissive)
MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance added by Syntology 2026-08 (from id) whyyyyy123/MemGuard/src/memguard/memory/governance.py 5f9acd760a957490 ran · our draft was wrong Apache-2.0 (permissive)
Beyond "AI Language": The case for the idiolectal nature of LLM output added by Syntology 2026-08 (from id) fsu-nlp/ai-idiolects/src/aiidiolects/compare_fingerprints.py e2efd9cc21506b21 ran fingerprinted MIT (permissive)
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents added by Syntology 2026-07 (from id) LYL1015/JarvisHub/apps/agents-cli/emp_code/scrape/component_card_index.py 722059afd24746a1 ran Apache-2.0 (permissive)
ReFact: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning added by Syntology 2026-07 (from id) NEUIR/REFACT/verl/verl/utils/reward_score/evdience_reward.py 81a9940880b8ee28 ran · our draft was wrong fingerprinted MIT (permissive)
Auditing Forgetting in Limited Memory Language Models added by Syntology 2026-07 (from id) raeesiarya/LMLMAudit/src/lmlm-audit/run_audit.py 56c931802f771660 ran · our draft was wrong fingerprinted MIT (permissive)
How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring added by Syntology 2026-06 (from id) mathieu0905/Code-Anchor/evaluation/compute_lexical_similarity.py 0ec01bca2dc31602 ran fingerprinted Apache-2.0 (permissive)
Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift added by Syntology 2026-06 (from id) kkr36/llm_detection/arxiv/inference_set_rewrite/analyze_ngram_overlap.py ebef45186043c5f8 ran fingerprinted no licence file found · pointer only
When Model Merging Breaks Routing: Training-Free Calibration for MoE added by Syntology 2026-06 (from id) huangcb01/HARC/src/merge_method/fisher.py 839be7ab6b2364ee ran no licence file found · pointer only
Poison with Style: A Practical Poisoning Attack on Code Large Language Models added by Syntology 2026-05 (from id) khangtran2020/pws/src/vllm_pred.py 5bc6e69d2f799430 ran no licence file found · pointer only
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval added by Syntology 2026-05 (from id) michalsr/ToolMerge/toolmerge/merging.py bd94326a00de8bb4 ran · our draft was wrong fingerprinted licence not identified · pointer only
The Neural Compiler: Program-to-Network Translation for Hybrid Scientific Machine Learning added by Syntology 2026-05 (from id) sheneman/neural_compiler/neural_compiler/parser/scheme_parser.py fe3db154d980e1cd ran fingerprinted MIT (permissive)
MasonNLP at MEDIQA-SYNUR 2026: Retrieval-Augmented Large Language Models for Schema-Constrained Clinical Information Extraction added by Syntology 2026-05 (from id) AHMRezaul/MEDIQA-SYNUR-2026/llama/rag.py d3c3053082b8a558 ran · our draft was wrong fingerprinted no licence file found · pointer only
MuonQ: Enhancing Low-Bit Muon Quantization via Directional Fidelity Optimization added by Syntology 2026-05 (from id) YupengSu/MuonQ/src/data_utils.py 55ccea02961066a2 ran Apache-2.0 (permissive)
LLM-Agnostic Semantic Representation Attack added by Syntology 2026-05 (from id) JiaweiLian/SRA/baselines/sra/utils.py ab33e2c8c310ea21 unverified MIT (permissive)
Topic Is Not Agenda: A Citation-Community Audit of Text Embeddings added by Syntology 2026-05 (from id) junseon-yoo/topic-not-agenda/code/eval/lexical_divergence.py ac9b1a0705cfa5b0 ran MIT (permissive)
SHARE: Social-Humanities AI for Research and Education added by Syntology 2026-04 (from id) allenai/s2_fos/src/s2_fos/training/open_ai_prompts.py 64b3241c045709a6 unverified Apache-2.0 (permissive)
Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional nonrobustness and how to fix them added by Syntology 2026-04 (from id) kohpangwei/group_DRO/dataset_scripts/generate_multinli.py 7ef9ae2541d2ae3e unverified MIT (permissive)
Universal Hypernetworks for Arbitrary Models added by Syntology 2026-04 (from id) Xuanfeng-Zhou/UHN/dataset/text_dataset.py 093adedba6fb6692 unverified MIT (permissive)
Optimsyn: Influence-Guided Rubrics Optimization for Synthetic Data Generation added by Syntology 2026-04 (from id) FanZT6/OptimSyn/training/influence/get_validation_dataset.py 5d4f58c63866080e unverified no licence file found · pointer only
Less Data, Faster Convergence: Goal-Driven Data Optimization for Multimodal Instruction Tuning added by Syntology 2026-03 (from id) rujiewu/GDO/gdo/extract_six_metrics.py 50bcada7381a2399 ran · our draft was wrong fingerprinted MIT (permissive)
The Chicken and Egg Dilemma: Co-optimizing Data and Model Configurations for LLMs added by Syntology 2026-02 (from id) princeton-nlp/LESS/less/data_selection/get_validation_dataset.py 5d4f58c63866080e unverified MIT (permissive)
IMU-1: Sample-Efficient Pre-training of Small Language Models added by Syntology 2026-02 (from id) thepowerfuldeez/sample_efficient_gpt/sample_efficient_gpt/evals/chat_web.py 4e22c143139d965f unverified no licence file found · pointer only
Enabling Progressive Whole-slide Image Analysis with Multi-scale Pyramidal Network added by Syntology 2026-02 (from id) mahmoodlab/CONCH/conch/open_clip_custom/custom_tokenizer.py 848bfe2419fbf48a unverified licence not identified · pointer only
Domain-Specific Knowledge Graphs in RAG-Enhanced Healthcare LLMs added by Syntology 2026-01 (from id) sydneyanuyah/RAGComparison/artifacts/csv_abstract_tokenizer.py 6d44e9d4930fe69a unverified no licence file found · pointer only
FINCARDS: Card-Based Analyst Reranking for Financial Document Question Answering added by Syntology 2026-01 (from id) XanderZhou2022/FINCARDS/pipeline/stage1_lexical_bm25.py 2ae5f9c473e7850b unverified no licence file found · pointer only
Far from the Shallow: Brain-Predictive Reasoning Embedding through Residual Disentanglement added by Syntology 2025-10 (from id) calclavia/tal-asrd/tal/alignment/aeneas.py d0e80b855a08c984 unverified no licence file found · pointer only
Automated Knowledge Graph Construction using Large Language Models and Sentence Complexity Modelling added by Syntology 2025-09 (from id) KaushikMahmud/CoDe-KG_EMNLP_2025/CoDe-KG/tokenize_paper.py f407977ea730a68c unverified MIT (permissive)
QFrCoLA: a Quebec-French Corpus of Linguistic Acceptability Judgments added by Syntology 2025-08 (from id) GRAAL-Research/QFrCoLA/article_src/dataset_analysis_cola_datasets.py ee8c55dae379f18f unverified licence not identified · pointer only
VidText: Towards Comprehensive Evaluation for Video Text Understanding 28 May 2025 shuyansy/vidtext/Evaluation/Evaluation.py 74cda7b9b94dbee6 ran · our draft was wrong fingerprinted no licence file found · pointer only
Synthetic Data RL: Task Definition Is All You Need 18 May 2025 gydpku/data_synthesis_rl/TinyZero/retriever.py 41f93bba51be41cf unverified Apache-2.0 (permissive)
JOLT-SQL: Joint Loss Tuning of Text-to-SQL with Confusion-aware Noisy Schema Sampling 20 May 2025 Songjw133/JOLT-SQL/test_suite_sql_eval/process_sql.py a906f5de1f5971b0 unverified Apache-2.0 (permissive)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation 2025-05 (from id) jiaqili3/DualCodec/dualcodec/dataset/processor.py 06cc7e2f6e744cda unverified MIT (permissive)
SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning 2025-04 (from id) IDEA-FinAI/SQL-R1/src/evaluations/spider1_evaluations/src/process_sql.py b72b76d185b39d84 unverified Apache-2.0 (permissive)
Primus: A Pioneering Collection of Open-Source Datasets for Cybersecurity LLM Training 16 Feb 2025 huggingface/cosmopedia/decontamination/decontaminate.py 9b49243ad0df4e21 unverified Apache-2.0 (permissive)
Rate of Model Collapse in Recursive Training 23 Dec 2024 berserank/rate-of-model-collapse/ngram_sim.py bc84e4715773bdc9 unverified no licence file found · pointer only
A Large-scale Empirical Study on Fine-tuning Large Language Models for Unit Testing 2024-12 (from id) iSEngLab/LLM4UT_Empirical/Finetune_Script/decoder_ag_main.py cadd56f46fae79b8 ran · our draft was wrong no licence file found · pointer only
Towards Agentic Schema Refinement 25 Nov 2024 agapiR/agentic-semantic-layer/src/process_sql.py a906f5de1f5971b0 unverified MIT (permissive)
AttriBoT: A Bag of Tricks for Efficiently Approximating Leave-One-Out Context Attribution 22 Nov 2024 r-three/AttriBoT/context_attribution/model_utils.py 0a40bf61a5a9fd27 unverified no licence file found · pointer only
Beyond Toxic Neurons: A Mechanistic Analysis of DPO for Toxicity Reduction 10 Nov 2024 yushi-y/dpo-toxic-neurons/evaluation/eval_utils.py 1a98f49d7c56acc9 unverified MIT (permissive)
TCP-Diffusion: A Multi-modal Diffusion Model for Global Tropical Cyclone Precipitation Forecasting with Change Awareness 17 Oct 2024 Zjut-MultimediaPlus/TCP-Diffusion/video_diffusion_pytorch/rainfall_diffusion_F4_E1ifs_0316.py 4e202ebb26added6 unverified no licence file found · pointer only
Rethinking Data Selection at Scale: Random Selection is Almost All You Need 12 Oct 2024 xiatingyu/sft-dataselection-at-scale/diverse/utils.py e8d1298ca322021e ran no licence file found · pointer only
Robust AI-Generated Text Detection by Restricted Embeddings 10 Oct 2024 silversolver/robustatd/fit_eraser_probing_tasks.py 6e90d22f1d7af677 unverified no licence file found · pointer only
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly 3 Oct 2024 princeton-nlp/helmet/model_utils.py 7257aebb0f11407a unverified MIT (permissive)
Procedure-Aware Surgical Video-language Pretraining with Hierarchical Knowledge Augmentation 30 Sep 2024 CAMMA-public/SurgVLP/surgvlp/surgvlp.py 2c5a6406461abcfa unverified no licence file found · pointer only
Improving Virtual Try-On with Garment-focused Diffusion Models 12 Sep 2024 siqi0905/gardiff/src/models/mask_attention.py 534b549a2186b672 ran no licence file found · pointer only
The Unreasonable Ineffectiveness of Nucleus Sampling on Mitigating Text Memorization 29 Aug 2024 lukaborec/memorization-nucleus-sampling/memorization/core/running_experiments.py 23c893fd630bb704 unverified no licence file found · pointer only
MSCPT: Few-shot Whole Slide Image Classification with Multi-scale and Context-focused Prompt Tuning 21 Aug 2024 hanminghao/mscpt/plot_heatmap.py 046182617bffcc33 ran no licence file found · pointer only
MAG-SQL: Multi-Agent Generative Approach with Soft Schema Linking and Iterative Sub-SQL Refinement for Text-to-SQL 15 Aug 2024 LancelotXWX/MAG-SQL/evaluation/process_sql.py a906f5de1f5971b0 unverified MIT (permissive)
CL-DiffPhyCon: Closed-loop Diffusion Control of Complex Physical Systems 31 Jul 2024 AI4Science-WestlakeU/CL_DiffPhyCon/model/text.py 03dfcc33eeaff851 unverified MIT (permissive)
Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism 24 Jul 2024 eit-nlp/2d-coordinate-system-for-icl/task_recognition_pir.py b9aa70482de54509 ran MIT (permissive)
TAGCOS: Task-agnostic Gradient Clustered Coreset Selection for Instruction Tuning Data 21 Jul 2024 2003pro/tagcos/data_selection/get_validation_dataset.py 5d4f58c63866080e unverified no licence file found · pointer only
Investigating Low-Rank Training in Transformer Language Models: Efficiency and Scaling Analysis 13 Jul 2024 CLAIRE-Labo/StructuredFFN/src/utils/refinedweb_llama.py 6c25271fe6f1f40f unverified no licence file found · pointer only
Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation 4 Jul 2024 LAMDA-RL/Q-Adapter/utils/datasets.py 6a216d37018e6479 ran Apache-2.0 (permissive)
BM25S: Orders of magnitude faster lexical search via eager sparse scoring 4 Jul 2024 xhluca/bm25s/bm25s/tokenization.py 7c70347879b998bf unverified MIT (permissive)
Understanding and Mitigating Language Confusion in LLMs 28 Jun 2024 for-ai/language-confusion/compute_metrics.py 3f7e1abc53ea4130 ran fingerprinted Apache-2.0 (permissive)
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs 18 Jun 2024 dakingrai/ood-generalization-semantic-boundary-techniques/evaluations/src/process_sql.py a906f5de1f5971b0 unverified MIT (permissive)
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs 11 Jun 2024 UKPLab/acl2024-DARA/utils/data_utils.py 2b51c2cec57bb498 ran no licence file found · pointer only
CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language Alignment 7 Jun 2024 iyyakuttiiyappan/CPLIP/zeroshot_classification.py 046182617bffcc33 ran no licence file found · pointer only
Harnessing the Power of MLLMs for Transferable Text-to-Image Person ReID 8 May 2024 wentaotan/mllm4text-reid/datasets/bases.py 3d862712396854e3 ran no licence file found · pointer only
RSCaMa: Remote Sensing Image Change Captioning with State Space Model 29 Apr 2024 chen-yang-liu/rscama/preprocess_data.py c6e897c47140c3e2 ran · our draft was wrong no licence file found · pointer only
In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter Optimization 25 Apr 2024 automl/ifbo/ifbo/surrogate.py d2413f6e81b6146b unverified MIT (permissive)
Continual Learning of Large Language Models: A Comprehensive Survey 25 Apr 2024 beyonderxx/trace/metrics.py e605721addd552ac ran · our draft was wrong fingerprinted Apache-2.0 (permissive)
PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks 6 Apr 2024 hrl-team/PhyloLM/lanlab/core/module/models/hf_models.py 02b77fb08ddcae25 ran GPL-3.0 (copyleft) · pointer only
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation 5 Apr 2024 tqzhong/cg4mctg/meta-mctg/contrastive_prefix_meta.py 9dbd25c2cddac479 ran MIT (permissive)
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation 5 Apr 2024 tqzhong/cg4mctg/meta-mctg/ctrl_meta.py 59ac6441acc40e2e ran MIT (permissive)
Benchmarking and Improving Compositional Generalization of Multi-aspect Controllable Text Generation 5 Apr 2024 tqzhong/cg4mctg/meta-mctg/dcg_meta.py 6c3ca0db0954ed64 ran MIT (permissive)
RSMamba: Remote Sensing Image Classification with State Space Model 28 Mar 2024 chen-yang-liu/change-agent/Multi_change/preprocess_data.py c6e897c47140c3e2 ran · our draft was wrong MIT (permissive)
Scaling Laws For Dense Retrieval 27 Mar 2024 jingtaozhan/drscale/dataset.py c8927edacebe2572 ran MIT (permissive)
SmallToLarge (S2L): Scalable Data Selection for Fine-tuning Large Language Models by Summarizing Training Trajectories of Small Models 12 Mar 2024 bigml-cs-ucla/s2l/utils.py e8d1298ca322021e ran MIT (permissive)
Analysis of Privacy Leakage in Federated Large Language Models 2 Mar 2024 vunhatminh/fl_attacks/LLMs/layers_attn_ldp.py 73b18550fdf42bdb unverified no licence file found · pointer only
Common 7B Language Models Already Possess Strong Math Capabilities 7 Mar 2024 jerrywu-code/susgen/src/finetune.py 448dfa384a102a35 unverified MIT (permissive)
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation 28 Feb 2024 xsc1234/info-rag/training/train_info_rag.py f49089050af8bfda unverified no licence file found · pointer only
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs 22 Feb 2024 aadityasingh/tokenizationcounts/utils.py 06f72a4bbe59fa49 unverified MIT (permissive)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text 17 Feb 2024 bohanhou14/semstamp/paraphrase_gen_utils.py eed1c0b474e01b4a ran no licence file found · pointer only
LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset 14 Feb 2024 osu-nlp-group/llm4chem/generation.py adcef27608cdbd15 unverified MIT (permissive)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 google-deepmind/searchless_chess/src/tokenizer.py 441d1203a85b9ca6 ran Apache-2.0 (permissive)
LESS: Selecting Influential Data for Targeted Instruction Tuning 6 Feb 2024 princeton-nlp/less/less/data_selection/get_validation_dataset.py 5d4f58c63866080e unverified MIT (permissive)
Inconsistent dialogue responses and how to recover from them 18 Jan 2024 mianzhang/cider/src/utils.py ac28a2d89a01c888 ran no licence file found · pointer only
A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity 3 Jan 2024 ajyl/dpo_toxic/toxicity/eval_interventions/eval_utils.py 9727f348a587c477 ran MIT (permissive)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering 3 Jan 2024 ByZ0e/Glance-Focus/dataset/nextqa.py c31fe80a1b2ba813 ran MIT (permissive)
RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair 25 Dec 2023 identical code first harvested elsewhere cadd56f46fae79b8 ran · our draft was wrong licence of this copy not recorded
RaDialog: A Large Vision-Language Model for Radiology Report Generation and Conversational Assistance 30 Nov 2023 chantalmp/radialog/chexbert/src/bert_tokenizer.py c98cb8be7d52c5c5 unverified no licence file found · pointer only
Boosting the Power of Small Multimodal Reasoning Models to Match Larger Models with Self-Consistency Training 23 Nov 2023 identical code first harvested elsewhere e605721addd552ac ran · our draft was wrong fingerprinted licence of this copy not recorded
Self-Evolved Diverse Data Sampling for Efficient Instruction Tuning 14 Nov 2023 ofa-sys/diverseevol/utils.py e8d1298ca322021e ran no licence file found · pointer only
Recognize Any Regions 2 Nov 2023 Surrey-UPLab/Recognize-Any-Regions/regionspot/modeling/clip/clip.py fda70a3344a4ff1d unverified licence not identified · pointer only
ACT-SQL: In-Context Learning for Text-to-SQL with Automatically-Generated Chain-of-Thought 26 Oct 2023 x-lance/text2sql-gpt/eval/process_sql.py a906f5de1f5971b0 unverified no licence file found · pointer only
H2O Open Ecosystem for State-of-the-art Large Language Models 17 Oct 2023 h2oai/h2ogpt/finetune.py 38cae2d1ced7906a ran Apache-2.0 (permissive)
Watermarking LLMs with Weight Quantization 17 Oct 2023 twilight92z/quantize-watermark/code/sft_data.py 79e4422d6f6bff4e unverified no licence file found · pointer only
Watermarking LLMs with Weight Quantization 17 Oct 2023 twilight92z/quantize-watermark/code/maintain_fp32/data_processor.py dd30351922015b11 unverified no licence file found · pointer only
EX-FEVER: A Dataset for Multi-hop Explainable Fact Verification 15 Oct 2023 identical code first harvested elsewhere 77535ec3fc559506 unverified licence of this copy not recorded
FinGPT: Instruction Tuning Benchmark for Open-Source Large Language Models in Financial Datasets 7 Oct 2023 AI4Finance-Foundation/FinGPT/fingpt/FinGPT_Benchmark/utils.py a8301984bfcfdc76 ran MIT (permissive)
LLaSM: Large Language and Speech Model 30 Aug 2023 linksoul-ai/llasm/infer_tokenize.py b2f9ced00c0053ba ran Apache-2.0 (permissive)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification 19 Aug 2023 QinYang79/RDE/2024-CVPR-RDE/datasets/bases.py 3d862712396854e3 ran no licence file found · pointer only
Enhancing Human-like Multi-Modal Reasoning: A New Challenging Dataset and Comprehensive Framework 24 Jul 2023 weijingxuan/COCO-MMR/evaluations.py e605721addd552ac ran · our draft was wrong fingerprinted no licence file found · pointer only
First-Explore, then Exploit: Meta-Learning to Solve Hard Exploration-Exploitation Trade-Offs 5 Jul 2023 btnorman/First-Explore/tiny_world/tiny_world_first_explore.py 6deaee162fc9967d ran · our draft was wrong MIT (permissive)
ProbVLM: Probabilistic Adapter for Frozen Vision-Language Models 1 Jul 2023 explainableml/probvlm/src/ds/_transforms.py 841979de2ba28e2f unverified MIT (permissive)
Interleaving Pre-Trained Language Models and Large Language Models for Zero-Shot NL2SQL Generation 15 Jun 2023 ruc-datalab/zeronl2sql/get_colval_map.py 2d9b155996a531f8 unverified MIT (permissive)
Linear Classifier: An Often-Forgotten Baseline for Text Classification 12 Jun 2023 ASUS-AICS/LibMultiLabel/libmultilabel/nn/data_utils.py 64d887c8cf8c806c unverified MIT (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections