Home › Code › load_vocab

load_vocab

Syntologyentry name in harvested coderead from the graph 2026-09-24

load_vocab appears in the code Syntology harvested for 107 papers, as 55 distinct code bodies found in 118 places (a place is one code body under one paper). At least one of them ran in 45 of the papers; 0 of the code bodies carry a behaviour fingerprint.

What this page is not. Routines are grouped here by the exact string of their function or class name. Nothing asserts that two samples named load_vocab do the same thing, share code, or are comparable; the name is a string, not an identity. Behaviour outputs (what a fingerprinted sample returned on the shared battery) are not in this export and are not shown here; the graph at syntology.ai holds them. "Ran" means executed on a synthesized fixture, not that the code is correct or reproduces a paper.

Samples Syntology

Syntology ran 14 of the 55 distinct code bodies named load_vocab; 41 are unverified. One tile per status, in the site's fixed vocabulary, each code body counted once:

1ran · honoured contract
0ran · violated contract
7ran · our draft was wrong
0ran · fixture could not drive it
6ran
41unverified
0fingerprinted

Licence is a property of each copy, so it is counted per place: 11 of the 118 places are pointer only (Syntology does not serve that copy's text). This site shows no code text for any sample; every row below links to the file in its repository where the record names one.

“Ran” means the sample executed on a synthesized input; it does not mean the output is correct. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code, and those samples did run. The ran count above is every status except unverified, the same rule as each paper page.

Papers

100 papers shown of 107 (newest first; the JSON twin carries all 118 places); 111 places in the table. A paper with no recorded date is placed by the month its arXiv id encodes, shown in the Date column as YYYY-MM (from id). One row per place: a paper whose repository defines the name more than once appears more than once, and the same code body held for several papers appears once under each, with the same status. Titles and dates are the archive's archive 2025-07-28 for papers in the archive, and the graph's for 2 papers added by Syntology; 12 papers have no page here and are shown by arXiv id only. Status and fingerprint are Syntology's record of each code body; licence is recorded for each place. The File cell ends with the code body's code_sha256, Syntology's identity for that exact code: an agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

PaperDateFileStatus SyntologyLicence
Federated generative event models for tokenized electronic health records added by Syntology 2026-08 (from id) bbj-lab/coreopsis/recipes/analyze-site-data.py d2c193ef453aefa3 unverified MIT (permissive)
Rethinking Drug-Drug Interaction Modeling as Generalizable Relation Learning added by Syntology 2026-01 (from id) SZU-ADDG/GenRel-DDI/model.py 8c1c03859eece728 ran · our draft was wrong no licence file found · pointer only
Unpacking Positional Encoding in Transformers: A Spectral Analysis of Content-Position Coupling 19 May 2025 Qihuai27/Deposit-Pattern-Research/dis_gen/1_data_preparation/data-gen.py 4f79bf4713e7fbec ran · our draft was wrong MIT (permissive)
MMedPO: Aligning Medical Vision-Language Models with Clinical-Aware Multimodal Preference Optimization 9 Dec 2024 aiming-lab/mmedpo/curation/Sample_Zero-Shot_Grounding_RSNA/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
AnyText2: Visual Text Generation and Editing With Customizable Attributes 22 Nov 2024 tyxsspa/anytext2/bert_tokenizer.py aaf1cb2c01043065 unverified Apache-2.0 (permissive)
VL-Uncertainty: Detecting Hallucination in Large Vision-Language Model via Uncertainty Estimation 18 Nov 2024 JT-Sun/Filtering-WoRA/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Semantic-Aligned Adversarial Evolution Triangle for High-Transferability Vision-Language Attack 4 Nov 2024 jiaxiaojunqaq/sa-aet/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 18 Sep 2024 MindCode-4/code-4/prophetnet/tokenization_prophetnet.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
PyMarian: Fast Neural Machine Translation and Evaluation in Python 15 Aug 2024 OpenNMT/CTranslate2/python/ctranslate2/converters/marian.py 241f78b9e78ef2ea ran MIT (permissive)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs 27 Jun 2024 RussianNLP/RuBLiMP/src/utils/data_loaders.py dbe2a14f4e641a48 ran Apache-2.0 (permissive)
One Perturbation is Enough: On Generating Universal Adversarial Perturbations against Vision-Language Pre-training Models 8 Jun 2024 ffhibnese/cpgc_vlp_universal_attacks/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong no licence file found · pointer only
Revisiting the Adversarial Robustness of Vision Language Models: a Multimodal Perspective 30 Apr 2024 ellezwq/mmcoa/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Self-Guided Masked Autoencoders for Domain-Agnostic Self-Supervised Learning 22 Feb 2024 johnathan-xie/sma/src/transformers/models/sma/tokenization_sma.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
UniChest: Conquer-and-Divide Pre-training for Multi-Source Chest X-Ray Classification 18 Dec 2023 elfenreigen/unichest/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Recursive Visual Programming 4 Dec 2023 para-lost/rvp/base_models/tcl/tcl_tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong no licence file found · pointer only
AnyText: Multilingual Visual Text Generation And Editing 6 Nov 2023 tyxsspa/anytext/bert_tokenizer.py aaf1cb2c01043065 unverified Apache-2.0 (permissive)
TLM: Token-Level Masking for Transformers 28 Oct 2023 blcuicall/CCL2022-CLTC/baselines/track2/tokenization.py deb4d1793cf3ba73 ran no licence file found · pointer only
ConPET: Continual Parameter-Efficient Tuning for Large Language Models 26 Sep 2023 raincleared-song/conpet/bmt_models/bee_tokenizer.py 8b48cf849a2af11c ran no licence file found · pointer only
Contrastive Grouping with Transformer for Referring Image Segmentation 2 Sep 2023 toneyaya/cgformer/bert/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages 23 Aug 2023 openbmb/viscpm/VisCPM/cpm_tokenizers/bee.py 83504f51b50ea9e6 ran no licence file found · pointer only
MESED: A Multi-modal Entity Set Expansion Dataset with Fine-grained Semantic Classes and Hard Negative Entities 27 Jul 2023 thukelab/mesed/utils.py 83911c412016bbe8 ran no licence file found · pointer only
Set-level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-training Models 26 Jul 2023 Zoky-2020/Set-level_Guidance_Attack/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Dataset for Pre-training and Benchmarks 7 Jun 2023 x-plug/youku-mplug/models/tokenization_mplug.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Towards Unified Text-based Person Retrieval: A Large-scale Multi-Attribute and Language Search Benchmark 5 Jun 2023 Shuyu-XJTU/APTM/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 29 May 2023 TXH-mercury/VALOR/model/bert_tokenizer.py fb95c4b13cdf89a2 unverified MIT (permissive)
RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search 23 May 2023 Flame-Chasers/RaSa/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed Videos 4 May 2023 MILVLG/anetqa-code/hcrn/DataLoader.py 4eb168a439e7658b unverified Apache-2.0 (permissive)
Dynamic Graph Enhanced Contrastive Learning for Chest X-ray Report Generation 18 Mar 2023 mlii0117/dcl/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
NapSS: Paragraph-level Medical Text Simplification via Narrative Prompting and Sentence-matching Summarization 11 Feb 2023 google-research/bert/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video 1 Feb 2023 X-PLUG/mPLUG-2/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Self-supervised vision-language pretraining for Medical visual question answering 24 Nov 2022 pengfeiliheu/m2i2/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
CRIPP-VQA: Counterfactual Reasoning about Implicit Physical Properties via Video Question Answering 7 Nov 2022 thaolmk54/hcrn-videoqa/DataLoader.py 4eb168a439e7658b unverified Apache-2.0 (permissive)
Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese 2 Nov 2022 ofa-sys/chinese-clip/cn_clip/clip/bert_tokenizer.py aaf1cb2c01043065 unverified MIT (permissive)
Towards Realistic Low-resource Relation Extraction: A Benchmark with Empirical Baseline Study 19 Oct 2022 zjunlp/KnowPrompt/models/bert/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative Grounding 11 Aug 2022 dzh19990407/ppmn/models/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training 12 Jul 2022 xinyu1205/idea-pytorch/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
PEVL: Position-enhanced Pre-training and Prompt Tuning for Vision-language Models 23 May 2022 thunlp/PEVL/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
EVA2.0: Investigating Open-Domain Chinese Dialogue Systems with Large-Scale Pre-Training 17 Mar 2022 thu-coai/EVA/src/tokenization_eva.py deb4d1793cf3ba73 ran MIT (permissive)
DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing 18 Nov 2021 microsoft/DeBERTa/DeBERTa/deberta/cache_utils.py 82c76e446399db0c unverified MIT (permissive)
Open Vocabulary Object Detection with Pseudo Bounding-Box Labels 18 Nov 2021 salesforce/pb-ovd/ALBEF/models/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong BSD-3-Clause (permissive)
The NiuTrans System for WNGT 2020 Efficiency Task 16 Sep 2021 NiuTrans/NiuTrans.NMT/tools/PrepareParallelData.py 5535dc12255780ab ran · honoured contract Apache-2.0 (permissive)
Graph Based Network with Contextualized Representations of Turns in Dialogue 9 Sep 2021 blacknoodle/tucore-gcn/models/BERT/tokenization.py 91bbcb76dbd7d974 unverified MIT (permissive)
WebQA: Multihop and Multimodal QA 1 Sep 2021 shubham-gupta-iitr/mmmlX/pytorch_pretrained_bert/tokenization.py 232295eeb04a142e unverified Apache-2.0 recorded; this copy not marked cleared · pointer only
BROS: A Pre-trained Language Model Focusing on Text and Layout for Better Key Information Extraction from Documents 10 Aug 2021 clovaai/bros/bros/tokenization_bros.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Curriculum learning for language modeling 4 Aug 2021 spacemanidol/CurriculumLearningForLanguageModels/bin/make_vocab_based_curriculum.py 637b815724dea79d unverified MIT (permissive)
CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark 15 Jun 2021 cbluebenchmark/cblue/cblue/models/zen/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
MathBERT: A Pre-trained Language Model for General NLP Tasks in Mathematics Education 2 Jun 2021 tbs17/MathBERT/mathbert/tokenization.py ff83ccc8b0b6462d unverified MIT recorded; this copy not marked cleared · pointer only
B-PROP: Bootstrapped Pre-training with Representative Words Prediction for Ad-hoc Retrieval 20 Apr 2021 Albert-Ma/PROP/pytorch_pretrain_bert/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
Designing a Minimal Retrieve-and-Read System for Open-Domain Question Answering 15 Apr 2021 clovaai/minimal-rnr-qa/playground/workspace/minimal_rnr/tfserving/bert_tokenizer/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Transformer Feed-Forward Layers Are Key-Value Memories 29 Dec 2020 mega002/ff-layers/analysis/key_value_agreement.py 906ed5859e0c2204 ran · our draft was wrong MIT (permissive)
Improving Multilingual Models with Language-Clustered Vocabularies 24 Oct 2020 afshinrahimi/mmner/config.py b1e144bdab6449a6 unverified Apache-2.0 (permissive)
TUTA: Tree-based Transformers for Generally Structured Table Pre-training 21 Oct 2020 microsoft/TUTA_table_understanding/tuta/tokenizer.py c6f97c063d30a371 unverified MIT (permissive)
Improving BERT Performance for Aspect-Based Sentiment Analysis 22 Oct 2020 IMPLabUniPr/BERT-for-ABSA/src/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
Deep Reinforcement Learning with Stacked Hierarchical Attention for Text-based Games 22 Oct 2020 YunqiuXu/SHA-KG/env.py ba9616834d32858b ran · our draft was wrong MIT (permissive)
Incorporating BERT into Parallel Sequence Decoding with Adapters 13 Oct 2020 lemmonation/abnet/bert/tokenization.py a4e183f24a5ad1a7 unverified MIT (permissive)
An Empirical Study on Large-Scale Multi-Label Text Classification Including Few and Zero-Shot Labels 4 Oct 2020 iliaschalkidis/lmtc-eurlex57k/neural_networks/layers/bert_tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding 11 Sep 2020 indobenchmark/indonlu/utils/functions.py 4384e3778f7c4a96 unverified Apache-2.0 (permissive)
The Lottery Ticket Hypothesis for Pre-trained BERT Networks 23 Jul 2020 TAMU-VITA/BERT-Tickets/transformers-master/src/transformers/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT (permissive)
Synthesize, Execute and Debug: Learning to Repair for Neural Program Synthesis 16 Jul 2020 sunblaze-ucb/SED/program_synthesis/datasets/data.py 3d825f33d60f08a3 unverified MIT (permissive)
DocVQA: A Dataset for VQA on Document Images 1 Jul 2020 anisha2102/docvqa/tokenization.py ff83ccc8b0b6462d unverified MIT (permissive)
How to Avoid Being Eaten by a Grue: Structured Exploration Strategies for Textual Worlds 12 Jun 2020 rajammanabrolu/Q-BERT/qbert/env.py ba9616834d32858b ran · our draft was wrong MIT (permissive)
On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines 8 Jun 2020 uds-lsv/bert-stable-fine-tuning/src/transformers/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong Apache-2.0 (permissive)
Exploring Cross-sentence Contexts for Named Entity Recognition with BERT 2 Jun 2020 jouniluoma/bert-ner-cmv/bert_tokenization.py 30b05f2a415dd132 unverified MIT (permissive)
Common Sense or World Knowledge? Investigating Adapter-Based Knowledge Injection into Pretrained Transformers 24 May 2020 wluper/retrograph/retrograph/modeling/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
Named Entity Recognition as Dependency Parsing 14 May 2020 juntaoy/biaffine-ner/extract_bert_features/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word Order 24 Apr 2020 huawei-noah/Pretrained-Language-Model/PMLM/interactive_conditional_samples_sincos_acrostic.py a87406334d83c277 unverified no licence file found · pointer only
MPNet: Masked and Permuted Pre-training for Language Understanding 20 Apr 2020 michael-wzhu/mpnet_zh/src/tokenization_bert.py e7fbc7a74a3457c7 ran · our draft was wrong MIT recorded; this copy not marked cleared · pointer only
The Right Tool for the Job: Matching Model and Instance Complexities 16 Apr 2020 allenai/sledgehammer/allennlp_overrides/pytorch_pretrained_bert/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
Generating Narrative Text in a Switching Dynamical System 8 Apr 2020 StonyBrookNLP/SLDS-Stories/data_utils.py a127d4e3ccc2f275 unverified MIT (permissive)
SentenceMIM: A Latent Variable Language Model 18 Feb 2020 seraphlabs-ca/SentenceMIM-demo/auxiliary.py 21afa721a457190d unverified MIT (permissive)
Interactive Refinement of Cross-Lingual Word Embeddings 8 Nov 2019 forest-snow/clime-ui/interface/load.py 357bd3b78248ba9a unverified MIT (permissive)
Do Multi-hop Readers Dream of Reasoning Chains? 31 Oct 2019 helloeve/bert-co-matching/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
Structured Pruning of Large Language Models 10 Oct 2019 Holldean/BERT-Pruning/bert/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 26 Sep 2019 kpe/bert-for-tf2/bert/tokenization/bert_tokenization.py 30b05f2a415dd132 unverified MIT (permissive)
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 26 Sep 2019 Soikonomou/bert_new/src/model/BERT/tokenization_bert.py 659f3761c1dc81fa unverified Apache-2.0 (permissive)
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models 23 Aug 2019 Arthurizijar/Bert_Airport/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models 23 Aug 2019 Maz101/Bert/tokenization.py 30b05f2a415dd132 unverified Apache-2.0 (permissive)
Well-Read Students Learn Better: On the Importance of Pre-training Compact Models 23 Aug 2019 coronazap/bert_client/tokenization.py 21c6f1a7ac31a2bd unverified Apache-2.0 (permissive)
Real-Time Open-Domain Question Answering with Dense-Sparse Phrase Index 13 Jun 2019 uwnlp/denspi/tokenization.py c8956ce86812a6c2 unverified Apache-2.0 (permissive)
Open Sesame: Getting Inside BERT's Linguistic Knowledge 4 Jun 2019 yongjie-lin/bert-opensesame/pytorch_pretrained_bert/tokenization.py fb95c4b13cdf89a2 unverified Apache-2.0 (permissive)
Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks 27 May 2019 NVIDIA/OpenSeq2Seq/ctc_decoder_with_lm/ctc-test.py 00c6b131afd785b5 unverified Apache-2.0 (permissive)
Deeper Text Understanding for IR with Contextual Neural Language Modeling 22 May 2019 AdeDZY/SIGIR19-BERT-IR/tokenization.py ff83ccc8b0b6462d unverified BSD-3-Clause (permissive)
Utilizing BERT for Aspect-Based Sentiment Analysis via Constructing Auxiliary Sentence 22 Mar 2019 HSLCY/ABSA-BERT-pair/tokenization.py deb4d1793cf3ba73 ran MIT (permissive)
BERT for Joint Intent Classification and Slot Filling 28 Feb 2019 mangushev/intent_slot/assistant/functions/intent_slot/preprocessor/tokenization.py deb4d1793cf3ba73 ran MIT (permissive)
Extracting Multiple-Relations in One-Pass with Pre-Trained Transformers 4 Feb 2019 helloeve/mre-in-one-pass/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
BioBERT: a pre-trained biomedical language representation model for biomedical text mining 25 Jan 2019 dmis-lab/bern/biobert_ner/tokenization.py ff83ccc8b0b6462d unverified BSD-2-Clause (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 IBM/MAX-Text-Sentiment-Classifier/core/bert/tokenization.py d0e80220bcbf4ae0 ran · our draft was wrong Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 1wy/bert/tokenization.py ff83ccc8b0b6462d unverified Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 lonePatient/Bert-Multi-Label-Text-Classification/pybert/model/albert/tokenization_bert.py 659f3761c1dc81fa unverified MIT (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 IBM/MAX-Question-Answering/core/tokenization.py 97438b663cee62b4 unverified Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 MaZhiyuanBUAA/bert-tf1.4.0/tokenization.py 85608a0993acae75 unverified Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 guoyaohua/BERT-Chinese-Annotation/tokenization.py cfdf5fac37771edf unverified Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 meizi1114/bert/tokenization.py 71806f38591c80b6 unverified Apache-2.0 (permissive)
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding 11 Oct 2018 tyxr/bert/tokenization.py b67ad78f40e86899 unverified Apache-2.0 (permissive)
NAPS: Natural Program Synthesis Dataset 6 Jul 2018 nearai/program_synthesis/program_synthesis/algolisp/dataset/data.py 711c950a814e87b1 unverified Apache-2.0 (permissive)
Relational recurrent neural networks 5 Jun 2018 cheonbok94/Pytorch-Relational-Recurrent-Neural-networks/utils/data.py 508cfb82d084f33d unverified MIT (permissive)
Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning 14 Mar 2018 davidmascharka/tbd-nets/utils/clevr.py ba3c072b123cf122 unverified MIT (permissive)
Efficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided Attention 24 Oct 2017 CSTR-Edinburgh/ophelia/data_load.py ed89c632018002ae unverified Apache-2.0 (permissive)
Attention Is All You Need 12 Jun 2017 Kyubyong/transformer/model.py eea349e2e1aa954c ran · our draft was wrong Apache-2.0 (permissive)
Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation 26 Sep 2016 microsoft/BlingFire/ldbsrc/bert_base_tok/tokenization.py ff83ccc8b0b6462d unverified MIT (permissive)
Neural Architectures for Named Entity Recognition 4 Mar 2016 IBM/MAX-Named-Entity-Tagger/core/utils.py a0562f49c954c62c unverified Apache-2.0 (permissive)
End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF 4 Mar 2016 monologg/korean-ner-pytorch/utils.py 7a25966763d63ede unverified Apache-2.0 (permissive)
Listen, Attend and Spell 5 Aug 2015 WindQAQ/listen-attend-and-spell/utils/vocab_utils.py 75776719b114cdb1 unverified Apache-2.0 (permissive)
The Ubuntu Dialogue Corpus: A Large Dataset for Research in Unstructured Multi-Turn Dialogue Systems 30 Jun 2015 dennybritz/chatbot-retrieval/models/helpers.py d90e90d3ea78789a unverified MIT (permissive)
Convolutional Neural Networks for Sentence Classification 25 Aug 2014 yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/tokenization.py fb95c4b13cdf89a2 unverified MIT (permissive)
Convolutional Neural Networks for Sentence Classification 25 Aug 2014 yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/tokenization_morp.py 90d3ee4851872058 unverified MIT (permissive)
arXiv:aaai_6453 siat-nlp/TransDG/src/utils/data_loader.py b2a79364e5c3e474 unverified MIT (permissive)
arXiv:aaai_32468 airsplay/LXMERT/src/lxrt/tokenization.py a4e183f24a5ad1a7 unverified MIT (permissive)
arXiv:aaai_27969 TianyuGoGO/XPNG/models/tokenization.py a4e183f24a5ad1a7 unverified Apache-2.0 (permissive)
arXiv:aaai_17659 frankaging/Quasi-Attention-ABSA/code/util/tokenization.py 97e01db27f6a6fd2 unverified MIT (permissive)
arXiv:2025.findings-emnlp.1385 bowen-upenn/ControlText/bert_tokenizer.py aaf1cb2c01043065 unverified Apache-2.0 (permissive)

This site shows no code text; each File cell links to the file on GitHub at the repository's current default branch, which may have changed since the harvest. "Pointer only" means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence cell for the reason. Per-sample records for a paper are on its paper page under "Code Syntology ran".

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections