Papers › Convolutional Neural Networks for Sentence Classification

Convolutional Neural Networks for Sentence Classification

25 Aug 2014EMNLP 2014 10arXiv:1408.5882archive 2025-07-28

Yoon Kim

We report on a series of experiments with convolutional neural networks (CNN) trained on top of pre-trained word vectors for sentence-level classification tasks. We show that a simple CNN with little hyperparameter tuning and static vectors achieves excellent results on multiple benchmarks. Learning task-specific vectors through fine-tuning offers further gains in performance. We additionally propose a simple modification to the architecture to allow for the use of both task-specific and static vectors. The CNN models discussed herein improve upon the state of the art on 4 out of 7 tasks, which include sentiment analysis and question classification.

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="1408.5882")

Code

Syntology Ran 19 of 77 code samples harvested from 20 repositories linked to this paper; 58 have no recorded run. Of those that ran: 6 ran · honoured contract; 1 ran · violated contract; 11 ran · our draft was wrong; 1 ran · fixture could not drive it.

By repository: community (archive-listed): 77 samples from 20 repositories, 19 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

118 repositories listed; official and paper-mentioned ones first.

Coda-s/BJTU_NLP_Practice mentioned on GitHubpytorch report
DataZwer/BaseTensorFlowModel mentioned on GitHubtf report
DataZwer/CNNTextClassification mentioned on GitHubtf report
DongjunLee/text-cnn-tensorflow mentioned on GitHubtf report
HuihuiChyan/BJTUNLP_Practice2020 mentioned on GitHubpytorch report
HuihuiChyan/BJTUNLP_Practice2021 mentioned on GitHubpytorch report
IndicoDataSolutions/finetune mentioned on GitHubtfMPL-2.0 report
JacobLau0513/676-MBTI mentioned on GitHubtf report
Jarvx/text-classification-pytorch mentioned on GitHubpytorch report
Jerryten/model_nlp mentioned on GitHubtf report
ManuelVs/NNForTextClassification mentioned on GitHubtfMIT report
ManuelVs/NeuralNetworks mentioned on GitHubtfMIT report
MayukhSobo/Conv4RNN mentioned on GitHubtf report
SeonbeomKim/TensorFlow-TextCNN mentioned on GitHubtf report
Shawn1993/cnn-text-classification-pytorch mentioned on GitHubpytorchApache-2.0 report
ShindongLee/Sentence_Classifier_CNN mentioned on GitHubpytorchMIT report
TobiasLee/Text-Classification mentioned on GitHubtf report
TsingZ0/PFL-Non-IID mentioned on GitHubpytorch report
adi2103/AML-CoVe mentioned on GitHubtf report
afrozloya/charrec mentioned on GitHubtfApache-2.0 report
andrewfwalters/w266-final-project mentioned on GitHubpytorchMIT report
attardi/CNN_sentence mentioned on GitHubtf report
bentrevett/pytorch-sentiment-analysis mentioned on GitHubpytorch report
bmcclannahan/NLP-Sentiment mentioned on GitHubpytorch report
bplank/teaching-dl4nlp mentioned on GitHub report
brightmart/ai_law mentioned on GitHubtf report
chiemenz/automl_vs_hyperdrive mentioned on GitHubtfMIT report
cmasch/cnn-text-classification mentioned on GitHubtf report
cove-adml/adml-anon mentioned on GitHubpytorch report
ddajing/multilayer-cnn-text-classification mentioned on GitHubtfApache-2.0 report
delldu/TextCNN mentioned on GitHubpytorchApache-2.0 report
dennybritz/cnn-text-classification-tf mentioned on GitHubtfApache-2.0 report
facebookresearch/pytext mentioned on GitHubpytorchNOASSERTION report
hanfeng108/Language-Detection mentioned on GitHubpytorch report
inspirehep/magpie mentioned on GitHubMIT report
kh-kim/simple-ntc mentioned on GitHubpytorch report
linguishi/chinese_sentiment mentioned on GitHubtf report
lrank/Linguistic_adversity mentioned on GitHubtf report
lrank/Robust-Representation mentioned on GitHubtf report
nvshrao/Pytorch-GLUE mentioned on GitHubpytorch report
paper-cat/Sentence-Classifications mentioned on GitHubtfMIT report
pipidog/CNLP mentioned on GitHubtf report
pipidog/QNLP mentioned on GitHubtf report
piyush2896/CNN-Text-Classifier mentioned on GitHubtfMIT report
plkumjorn/FIND mentioned on GitHubtfGPL-3.0 report
prakashpandey9/Text-Classification-Pytorch mentioned on GitHubpytorchMIT report
pranjalg96/Stylized-Image-captioning mentioned on GitHubpytorch report
rainorangelemon/TextCNN-with-Attention mentioned on GitHubpytorchApache-2.0 report
rakeshmp18/Text_Analytics mentioned on GitHub report
reneang17/authorencoder mentioned on GitHubpytorch report
richinkabra/CoVe-BCN mentioned on GitHubpytorch report
thisisiron/kaggle-toxic-comment mentioned on GitHubtf report
threelittlemonkeys/text-cnn-pytorch mentioned on GitHubpytorch report
timerstime/SDG4DA mentioned on GitHubtf report
toru34/kim_emnlp_2014 mentioned on GitHub report
unik00/textCNN-pytorch mentioned on GitHubpytorch report
unik00/textCNN-tf2 mentioned on GitHubpytorch report
usama6832/attention_model mentioned on GitHub report
wiseodd/controlled-text-generation mentioned on GitHubpytorch report
wy-ei/Text-CNN mentioned on GitHubpytorchMIT report
xiaoxinyu1997/Knowledge-Graph mentioned on GitHubpytorch report
yangyucheng000/textcnn mentioned on GitHubmindspore report
yinghao1019/NLP_and_DL_practice mentioned on GitHubpytorch report
yoonkim/CNN_sentence mentioned on GitHubtf report
yschoi-nisp/AI-Grand-Challenge-2020 mentioned on GitHubpytorchMIT report
dmlc/gluon-nlp mxnetApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

77 samples harvested; 19 ran; 6 honoured the contract we drafted; 58 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

6ran · honoured contract
1ran · violated contract
11ran · our draft was wrong
1ran · fixture could not drive it
58unverified

Licence: 15 of the 77 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from 20 repositories linked to this paper, official or community; each sample names its own and says which. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

swish yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/modeling.py community (archive-listed) ran · our draft was wrong fingerprinted MIT (permissive) · 0f786c407fb1ee4c · report
clean_str audqhsid/-Review-CNN-for-Sentence-Classification/CNNforSentence/data_helpers.py community (archive-listed) ran · our draft was wrong fingerprinted no licence file found · pointer only · efe1a86fd9a2468e · report
compute_ratio svjan5/CNN-for-text-classification/SVM/SVM-1ofVEncoding.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 1f0857ef07d36305 · report
cost_weight_for_imbalanced_label SeonbeomKim/TensorFlow-TextCNN/SST_train.py community (archive-listed) ran · fixture could not drive it no licence file found · pointer only · 1826f73b65b1976e · report
gelu yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/modeling.py community (archive-listed) ran · honoured contract fingerprinted MIT (permissive) · 40e9fee2e0b7e278 · report
get_W usama6832/attention_model/process_sst1_data.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 8e159fcba115372b · report
get_dict Nirvanabc/sentence_classification/prepare_data.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 74bca835aa136142 · report
get_idx_from_sent usama6832/attention_model/mr_data.py community (archive-listed) ran · honoured contract no licence file found · pointer only · 2777a9ed50134acd · report
get_time_dif gaussic/text-classification-cnn-rnn/run_cnn.py community (archive-listed) ran · our draft was wrong MIT (permissive) · d9e00c99bc3b43d2 · report
load_data_and_labels audqhsid/-Review-CNN-for-Sentence-Classification/CNNforSentence/data_helpers.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 24886408b62be99a · report
make_idx_data_cv usama6832/attention_model/mr_data.py community (archive-listed) ran · honoured contract no licence file found · pointer only · c3873605307a3c77 · report
make_idx_data_cv_org_text usama6832/attention_model/mr_data.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 781a7c253f13fdb1 · report
normalize Nirvanabc/sentence_classification/prepare_data.py community (archive-listed) ran · violated contract fingerprinted no licence file found · pointer only · e55c5e84d71ce43a · report
read_text kh-kim/simple-ntc/classify.py community (archive-listed) ran · honoured contract fingerprinted no licence file found · pointer only · aa2e11b6e900e44a · report
read_word_and_its_vec Nirvanabc/sentence_classification/prepare_data.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 81912d07d41ad566 · report
tokenize svjan5/CNN-for-text-classification/SVM/SVM-1ofVEncoding.py community (archive-listed) ran · our draft was wrong no licence file found · pointer only · 79c1b7c0f8e87a6b · report
url_to_filename yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/file_utils.py community (archive-listed) ran · our draft was wrong fingerprinted MIT (permissive) · af64ec220e8bcdbc · report
warmup_constant yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/optimization.py community (archive-listed) ran · honoured contract fingerprinted MIT (permissive) · e7d542062316094a · report
warmup_linear yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/optimization.py community (archive-listed) ran · honoured contract fingerprinted MIT (permissive) · c58d57224530d17e · report
assert_and_compile_args piyush2896/CNN-Text-Classifier/util.py community (archive-listed) unverified MIT (permissive) · e46bdc997c311e22 · report
batches_generator piyush2896/CNN-Text-Classifier/nn/preprocessor.py community (archive-listed) unverified MIT (permissive) · c62648a11b00a100 · report
binary_accuracy piyush2896/CNN-Text-Classifier/nn/metrics.py community (archive-listed) unverified MIT (permissive) · d51abac955ff517c · report
binary_entropy piyush2896/CNN-Text-Classifier/nn/losses.py community (archive-listed) unverified MIT (permissive) · d89fbd714915f9e7 · report
build_dict wy-ei/Text-CNN/data.py community (archive-listed) unverified MIT (permissive) · 31bacc027a1637a7 · report
build_vocab alexander-rakhlin/CNN-for-Sentence-Classification-in-Keras/data_helpers.py community (archive-listed) unverified MIT (permissive) · b32b5833493d607a · report
cached_path yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/file_utils.py community (archive-listed) unverified MIT (permissive) · 349d780dd4a89a37 · report
clean_review rainorangelemon/TextCNN-with-Attention/clean.py community (archive-listed) unverified Apache-2.0 (permissive) · 6e49b503ffe317d7 · report
clean_str Shawn1993/cnn-text-classification-pytorch/mydatasets.py community (archive-listed) unverified Apache-2.0 (permissive) · 8ebb1beb61837fb1 · report
clean_str ddajing/multilayer-cnn-text-classification/data_helpers.py community (archive-listed) unverified Apache-2.0 (permissive) · efc3a77905da7a9d · report
collate_fn Shawn1993/cnn-text-classification-pytorch/mydatasets.py community (archive-listed) unverified Apache-2.0 (permissive) · dfd16b852a1f08c1 · report
compute_word2vec_for_phrase inspirehep/magpie/magpie/base/word2vec.py community (archive-listed) unverified MIT (permissive) · 7b9076aa00b98231 · report
concat piyush2896/CNN-Text-Classifier/nn/layers.py community (archive-listed) unverified MIT (permissive) · 4df1c013c4cd5872 · report
convert_examples_to_tokens yschoi-nisp/AI-Grand-Challenge-2020/BERT_total.py community (archive-listed) unverified MIT (permissive) · e7ecdb123ea7ddd2 · report
convert_tokens_to_features yschoi-nisp/AI-Grand-Challenge-2020/BERT_total.py community (archive-listed) unverified MIT (permissive) · 728fbd506ea36d0d · report
convert_tokens_to_features_eval yschoi-nisp/AI-Grand-Challenge-2020/BERT_total.py community (archive-listed) unverified MIT (permissive) · 9cf15995db9f9f06 · report
create_embedding_index piyush2896/CNN-Text-Classifier/util.py community (archive-listed) unverified MIT (permissive) · 26dfc0d2a5e7b9e7 · report
create_mask rainorangelemon/TextCNN-with-Attention/model.py community (archive-listed) unverified Apache-2.0 (permissive) · 3c77d69fa7310478 · report
dropout piyush2896/CNN-Text-Classifier/nn/layers.py community (archive-listed) unverified MIT (permissive) · 09cb65a28eb46830 · report
eval delldu/TextCNN/model.py community (archive-listed) unverified Apache-2.0 (permissive) · dd22129a171b50fb · report
evaluate wy-ei/Text-CNN/trainer.py community (archive-listed) unverified MIT (permissive) · 9a0485a6887c7747 · report
explained_variance piyush2896/CNN-Text-Classifier/nn/metrics.py community (archive-listed) unverified MIT (permissive) · c108960f22fe26aa · report
feed_data gaussic/text-classification-cnn-rnn/run_cnn.py community (archive-listed) unverified MIT (permissive) · 4515828165469544 · report
filename_to_url yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/file_utils.py community (archive-listed) unverified MIT (permissive) · db0ac56aaf6e35e6 · report
get_all_answers inspirehep/magpie/magpie/utils.py community (archive-listed) unverified MIT (permissive) · 8dc8d3a14cde192f · report
get_clones rainorangelemon/TextCNN-with-Attention/model.py community (archive-listed) unverified Apache-2.0 (permissive) · 5cb79ca7ceefc156 · report
get_polarity_dictionary chiemenz/AzureML-Sentiment-Classification-and-Model-Deployment/rating_ml_modules/featurization/word_sentiment_polarity/sentiment_polarity.py community (archive-listed) unverified MIT (permissive) · 99c0dc46fd15bb10 · report
input piyush2896/CNN-Text-Classifier/nn/layers.py community (archive-listed) unverified MIT (permissive) · de36f9c1e251692d · report
label_token delldu/TextCNN/data.py community (archive-listed) unverified Apache-2.0 (permissive) · 9e8508cb98e5104b · report
linear TobiasLee/Text-Classification/models/cnn.py community (archive-listed) unverified Apache-2.0 (permissive) · c970f81582bb1f97 · report
load_bin_vec usama6832/attention_model/process_sst1_data.py community (archive-listed) unverified no licence file found · pointer only · dc39e2e56c9504b1 · report
load_data ddajing/multilayer-cnn-text-classification/data_helpers.py community (archive-listed) unverified Apache-2.0 (permissive) · a59a82e26e9459f2 · report
load_dataset piyush2896/CNN-Text-Classifier/util.py community (archive-listed) unverified MIT (permissive) · 506f16da45fa7bd5 · report
load_from_disk inspirehep/magpie/magpie/utils.py community (archive-listed) unverified MIT (permissive) · 2c9841669c277781 · report
load_vocab yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/tokenization.py community (archive-listed) unverified MIT (permissive) · fb95c4b13cdf89a2 · report
load_vocab yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/tokenization_morp.py community (archive-listed) unverified MIT (permissive) · 90d3ee4851872058 · report
load_word2vec ddajing/multilayer-cnn-text-classification/data_helpers.py community (archive-listed) unverified Apache-2.0 (permissive) · 3f3f92f4ce36b0a7 · report
longest_sentence ShindongLee/Sentence_Classifier_CNN/preprocess.py community (archive-listed) unverified MIT (permissive) · 795300f6734869ed · report
mse piyush2896/CNN-Text-Classifier/nn/losses.py community (archive-listed) unverified MIT (permissive) · 8197ee6236e68110 · report
multiclass_accuracy piyush2896/CNN-Text-Classifier/nn/metrics.py community (archive-listed) unverified MIT (permissive) · bdb736f7b92a0ad3 · report
pad_sentences alexander-rakhlin/CNN-for-Sentence-Classification-in-Keras/data_helpers.py community (archive-listed) unverified MIT (permissive) · 8f5471ca03ed0239 · report
predict delldu/TextCNN/model.py community (archive-listed) unverified Apache-2.0 (permissive) · 60169e8e6c263250 · report
predict wy-ei/Text-CNN/trainer.py community (archive-listed) unverified MIT (permissive) · c5beea5bfbce4121 · report
read_glove_vectors shagunsodhani/CNN-Sentence-Classifier/app/reader/filereader.py community (archive-listed) unverified MIT (permissive) · 8dc6cef357628e1b · report
read_input_data shagunsodhani/CNN-Sentence-Classifier/app/reader/filereader.py community (archive-listed) unverified MIT (permissive) · 0402070ef0b2c2b3 · report
read_su_sentiment_rotten_tomatoes usama6832/attention_model/process_sst1_data.py community (archive-listed) unverified no licence file found · pointer only · 85763dbc456f9a11 · report
relabel_reviews chiemenz/AzureML-Sentiment-Classification-and-Model-Deployment/rating_ml_modules/featurization/preprocessing.py community (archive-listed) unverified MIT (permissive) · 4677d82ba2d19519 · report
softmax_entropy piyush2896/CNN-Text-Classifier/nn/losses.py community (archive-listed) unverified MIT (permissive) · bd97236dd26fd34d · report
text_token delldu/TextCNN/data.py community (archive-listed) unverified Apache-2.0 (permissive) · 55365320fd689119 · report
topic_dictionary_2_dataframe chiemenz/AzureML-Sentiment-Classification-and-Model-Deployment/rating_ml_modules/featurization/topic_modeling/topic_model.py community (archive-listed) unverified MIT (permissive) · 4f43bc8ef9323fc3 · report
train delldu/TextCNN/model.py community (archive-listed) unverified Apache-2.0 (permissive) · 619063371b6abd03 · report
train rainorangelemon/TextCNN-with-Attention/model.py community (archive-listed) unverified Apache-2.0 (permissive) · 059bfdf0c683e8f4 · report
train wy-ei/Text-CNN/trainer.py community (archive-listed) unverified MIT (permissive) · 346efb91ac88dfd5 · report
train_data_ready ShindongLee/Sentence_Classifier_CNN/preprocess.py community (archive-listed) unverified MIT (permissive) · e479b1df85a99a82 · report
train_test_split piyush2896/CNN-Text-Classifier/nn/preprocessor.py community (archive-listed) unverified MIT (permissive) · 9239e6af1bbb6e07 · report
validate ShindongLee/Sentence_Classifier_CNN/validation.py community (archive-listed) unverified MIT (permissive) · 3327aa7527520a84 · report
warmup_cosine yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/optimization.py community (archive-listed) unverified MIT (permissive) · 35f7cddf90dd05d4 · report
whitespace_tokenize yschoi-nisp/AI-Grand-Challenge-2020/pytorch_pretrained_bert/tokenization.py community (archive-listed) unverified MIT (permissive) · da7295883cd7da14 · report

Tasks

Emotion Recognition in ConversationGeneral ClassificationSentenceSentence ClassificationSentiment Analysis

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Emotion Recognition in Conversation CPED TextCNN Accuracy of Sentiment 48.90 #6 of 11 Archive leaderboard report
Emotion Recognition in Conversation CPED TextCNN Macro-F1 of Sentiment 34.37 #6 of 11 Archive leaderboard report
Sentiment Analysis SST-2 Binary classification CNN-multichannel [kim2013] Accuracy 88.1 #67 of 87 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections