Browse State-of-the-Art › Natural Language Processing
Natural Language Processing
Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 2,998 tables with Natural Language Processing; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).
Parent tasks
190 tasks in Natural Language Processing have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.
- Question Answering (19)
- Image Generation (23)
- Machine Translation (10)
- Link Prediction (6)
- Visual Question Answering (VQA) (9)
- Named Entity Recognition (NER) (13)
- Text Classification (16)
- Classification (24)
- Language Modelling (6)
- Relation Extraction (15)
- Sentiment Analysis (12)
- Text Summarization (12)
- Continual Learning (5)
- Natural Language Inference (3)
- Image Captioning (8)
- Visual Question Answering (5)
- Few-Shot Learning (12)
- Code Generation (5)
- Data-to-Text Generation (3)
- Common Sense Reasoning (17)
- Stance Detection (4)
- Document Classification (1)
- Text Generation (25)
- Semantic Parsing (7)
- Intent Detection (1)
- Aspect-Based Sentiment Analysis (ABSA) (8)
- Text-to-Image Generation (7)
- Video Generation (2)
- Prompt Engineering (1)
- Dependency Parsing (5)
- Coreference Resolution (2)
- Word Sense Disambiguation (1)
- Abstractive Text Summarization (3)
- Part-Of-Speech Tagging (1)
- Hate Speech Detection (3)
- Slot Filling (2)
- Semantic Textual Similarity (2)
- Dialogue Generation (1)
- Cross-Modal Retrieval (5)
- Grammatical Error Correction (1)
- Open Information Extraction (1)
- Retrieval (3)
- Information Retrieval (6)
- Question Generation (1)
- Speech-to-Text Translation (1)
- Entity Resolution (1)
- 2D Semantic Segmentation (17)
- Mathematical Reasoning (7)
- Text-To-SQL (1)
- Entity Alignment (1)
- Cross-Lingual Document Classification (1)
- Emotion Recognition (12)
- Event Extraction (3)
- Source Code Summarization (1)
- Relation Classification (3)
- Entity Typing (1)
- Few-Shot Text Classification (1)
- Table-to-Text Generation (1)
- Knowledge Distillation (2)
- Reading Comprehension (20)
- Semantic Similarity (2)
- Knowledge Graph Completion (4)
- Document Summarization (1)
- Semantic Role Labeling (3)
- Binary text classification (1)
- Natural Language Understanding (2)
- Optical Character Recognition (OCR) (10)
- Topic Models (2)
- Intrusion Detection (1)
- Language Identification (2)
- Sentence Classification (1)
- Text-To-Speech Synthesis (2)
- Text-to-Video Generation (2)
- Key Information Extraction (1)
- Aspect Extraction (2)
- Representation Learning (14)
- Deep Clustering (3)
- Story Generation (1)
- Bias Detection (1)
- Phrase Grounding (1)
- Extractive Text Summarization (1)
- Binary Classification (6)
- Task-Oriented Dialogue Systems (1)
- Constituency Parsing (1)
- Document-level Relation Extraction (1)
- Discourse Parsing (3)
- Document Text Classification (3)
- Data Augmentation (2)
- Paraphrase Generation (1)
- 3D Action Recognition (6)
- Multimodal Machine Translation (2)
- Text Clustering (3)
- Meme Classification (1)
- Age And Gender Classification (1)
- Inductive knowledge graph completion (1)
- Biomedical Information Retrieval (2)
- Multimodal Text and Image Classification (2)
- Text Style Transfer (3)
- Document Ranking (1)
- Term Extraction (2)
- Data-free Knowledge Distillation (1)
- Weakly Supervised Classification (1)
- Temporal Information Extraction (1)
- Hope Speech Detection (3)
- Handwriting Verification (1)
- Aspect-Category-Opinion-Sentiment Quadruple Extraction (1)
- Recognizing Emotion Cause in Conversations (1)
- Reinforcement Learning (1)
- Active Learning (1)
- Instruction Following (1)
- Cross-Lingual Transfer (2)
- Chatbot (1)
- Natural Language Queries (1)
- Argument Mining (6)
- Multimodal Deep Learning (1)
- Language Acquisition (1)
- Conversational Question Answering (1)
- Sentence Completion (1)
- Spam detection (2)
- Visual Storytelling (1)
- Conditional Text Generation (2)
- Open-Domain Dialog (1)
- Temporal Relation Extraction (1)
- Ad-Hoc Information Retrieval (1)
- Document AI (1)
- Goal-Oriented Dialog (1)
- Sentence Compression (1)
- Emotional Intelligence (3)
- Question Similarity (1)
- Lexical Normalization (1)
- Attribute Extraction (1)
- Clinical Concept Extraction (1)
- Scientific Document Summarization (1)
- Summarization (2)
- 1 Image, 2*2 Stitchi (12)
- Information Extraction (16)
- Text2text Generation (3)
- Dialogue (21)
- Language Modeling (1)
- Deep Learning (1)
- Large Language Model (3)
- Word Embeddings (4)
- NMT (1)
- Text to Image Generation (1)
- Sentence Embeddings (4)
- Symbolic Regression (1)
- document understanding (1)
- Data Integration (3)
- Model Editing (1)
- Authorship Attribution (1)
- De-identification (2)
- Spelling Correction (1)
- Token Classification (2)
- Dialogue Understanding (2)
- Abuse Detection (1)
- Constrained Clustering (2)
- Stock Prediction (3)
- Table annotation (6)
- Sentence Summarization (1)
- Conversational Response Generation (1)
- Propaganda detection (2)
- Twitter Sentiment Analysis (1)
- News Generation (1)
- Negation Detection (1)
- Cross-Lingual Entity Linking (1)
- Lexical Analysis (1)
- Literature Mining (1)
- Persian Sentiment Analysis (1)
- Sentence Pair Modeling (1)
- Data Mining (7)
- Continual Named Entity Recognition (1)
- Multimodal Association (1)
- Speculation Detection (1)
- Joint Multilingual Sentence Representations (1)
- Natural Language Transduction (1)
- Personality Generation (1)
- Political evalutation (1)
- trustable and focussed LLM generated content (1)
- Vietnamese Aspect-Based Sentiment Analysis (1)
- Vietnamese Sentiment Analysis (1)
- Nested Term Recognition (1)
- NLP based Person Retrival (1)
- Anaphora Resolution (2)
- Chinese (5)
- Cross-Lingual (4)
- Optical Charater Recogntion (1)
- Pcl Detection (2)
- Shallow Syntax (1)
- Taxonomy Learning (2)
- Temporal Processing (3)
Question Answering
142 benchmarks · 4,171 papers with codeMultiple Choice Question Answering (MCQA)
31 benchmarks · 37 papers with code
Zero-Shot Video Question Answer
17 benchmarks · 73 papers with code
Open-Domain Question Answering
15 benchmarks · 238 papers with code
Knowledge Base Question Answering
10 benchmarks · 67 papers with code
Answer Selection
6 benchmarks · 52 papers with code
5 shown of 19 sub-tasks (1 filed under another area). All sub-tasks of Question Answering →
Image Generation
93 benchmarks · 3,102 papers with codeImage-to-Image Translation
38 benchmarks · 550 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Image Inpainting
12 benchmarks · 331 papers with code
Conditional Image Generation
11 benchmarks · 167 papers with code
Layout-to-Image Generation
11 benchmarks · 24 papers with code
5 shown of 23 sub-tasks (6 filed under another area). All sub-tasks of Image Generation →
Machine Translation
84 benchmarks · 2,444 papers with codeUnsupervised Machine Translation
9 benchmarks · 33 papers with code
Multimodal Machine Translation
3 benchmarks · 38 papers with code
Low-Resource Neural Machine Translation
1 benchmark · 28 papers with code
Transliteration
0 benchmarks · 54 papers with code
Bilingual Lexicon Induction
0 benchmarks · 35 papers with code
5 shown of 10 sub-tasks (1 filed under another area). All sub-tasks of Machine Translation →
Link Prediction
80 benchmarks · 974 papers with codeDynamic Link Prediction
10 benchmarks · 20 papers with code
Inductive Link Prediction
3 benchmarks · 23 papers with code
Link prediction on DH-KGs
1 benchmark · 1 paper with code
Hyperedge Prediction
0 benchmarks · 10 papers with code
Anchor link prediction
0 benchmarks · 1 paper with code
5 shown of 6 sub-tasks (1 filed under another area). All sub-tasks of Link Prediction →
Visual Question Answering (VQA)
76 benchmarks · 1,039 papers with codeVisual Question Answering
29 benchmarks · 1,042 papers with code
Machine Reading Comprehension
4 benchmarks · 206 papers with code
Chart Question Answering
3 benchmarks · 25 papers with code
3D Question Answering (3D-QA)
3 benchmarks · 17 papers with code
Generative Visual Question Answering
1 benchmark · 7 papers with code
5 shown of 9 sub-tasks (2 filed under another area). All sub-tasks of Visual Question Answering (VQA) →
Named Entity Recognition (NER)
76 benchmarks · 955 papers with codeChinese Named Entity Recognition
7 benchmarks · 40 papers with code
Nested Named Entity Recognition
6 benchmarks · 46 papers with code
Multi-modal Named Entity Recognition
5 benchmarks · 6 papers with code
NER
4 benchmarks · 665 papers with code
Few-shot NER
4 benchmarks · 42 papers with code
5 shown of 13 sub-tasks. All sub-tasks of Named Entity Recognition (NER) →
Text Classification
68 benchmarks · 1,308 papers with codeDocument Classification
21 benchmarks · 235 papers with code
Multi-Label Text Classification
20 benchmarks · 78 papers with code
Emotion Classification
9 benchmarks · 117 papers with code
Few-Shot Text Classification
8 benchmarks · 46 papers with code
Topic Models
6 benchmarks · 229 papers with code
5 shown of 16 sub-tasks. All sub-tasks of Text Classification →
Classification
58 benchmarks · 3,778 papers with codeGraph Classification
73 benchmarks · 483 papers with code
Text Classification
68 benchmarks · 1,308 papers with code
Audio Classification
22 benchmarks · 183 papers with code
Medical Image Classification
11 benchmarks · 183 papers with code
Multi-class Classification
5 benchmarks · 289 papers with code
5 shown of 24 sub-tasks (5 filed under another area). All sub-tasks of Classification →
Language Modelling
55 benchmarks · 7,012 papers with codeLong-range modeling
2 benchmarks · 69 papers with code
Cross-Document Language Modeling
2 benchmarks · 1 paper with code
Protein Language Model
1 benchmark · 47 papers with code
XLM-R
0 benchmarks · 99 papers with code
Sentence Pair Modeling
0 benchmarks · 6 papers with code
5 shown of 6 sub-tasks. All sub-tasks of Language Modelling →
Relation Extraction
50 benchmarks · 735 papers with codeJoint Entity and Relation Extraction
16 benchmarks · 56 papers with code
Relation Classification
8 benchmarks · 160 papers with code
Document-level Relation Extraction
4 benchmarks · 66 papers with code
Dialog Relation Extraction
2 benchmarks · 13 papers with code
Relationship Extraction (Distant Supervised)
2 benchmarks · 11 papers with code
5 shown of 15 sub-tasks. All sub-tasks of Relation Extraction →
Sentiment Analysis
42 benchmarks · 1,509 papers with codeAspect-Based Sentiment Analysis (ABSA)
18 benchmarks · 185 papers with code
Multimodal Sentiment Analysis
5 benchmarks · 93 papers with code
Aspect Sentiment Triplet Extraction
4 benchmarks · 29 papers with code
Aspect Term Extraction and Sentiment Classification
1 benchmark · 8 papers with code
Arabic Sentiment Analysis
1 benchmark · 7 papers with code
5 shown of 12 sub-tasks. All sub-tasks of Sentiment Analysis →
Text Summarization
37 benchmarks · 440 papers with codeAbstractive Text Summarization
15 benchmarks · 362 papers with code
Document Summarization
7 benchmarks · 226 papers with code
Multi-Document Summarization
5 benchmarks · 113 papers with code
Extractive Text Summarization
5 benchmarks · 34 papers with code
Long-Form Narrative Summarization
4 benchmarks · 4 papers with code
5 shown of 12 sub-tasks. All sub-tasks of Text Summarization →
Continual Learning
33 benchmarks · 1,142 papers with codeClass Incremental Learning
6 benchmarks · 296 papers with code
TiROD
1 benchmark · 1 paper with code
Continual Named Entity Recognition
0 benchmarks · 3 papers with code
Continual Panoptic Segmentation
0 benchmarks · 2 papers with code
unsupervised class-incremental learning
0 benchmarks · 2 papers with code
5 shown of 5 sub-tasks (1 filed under another area).
Natural Language Inference
33 benchmarks · 821 papers with codeCross-Lingual Natural Language Inference
4 benchmarks · 17 papers with code
Visual Entailment
3 benchmarks · 33 papers with code
Answer Generation
2 benchmarks · 111 papers with code
3 shown of 3 sub-tasks.
Image Captioning
33 benchmarks · 774 papers with codeSemi Supervised Learning for Image Captioning
3 benchmarks · 2 papers with code
3D dense captioning
2 benchmarks · 13 papers with code
Hindi Image Captioning
2 benchmarks · 0 papers with code
Relational Captioning
1 benchmark · 2 papers with code
controllable image captioning
0 benchmarks · 8 papers with code
5 shown of 8 sub-tasks (1 filed under another area). All sub-tasks of Image Captioning →
Visual Question Answering
29 benchmarks · 1,042 papers with codeSpatial Reasoning
2 benchmarks · 198 papers with code
Explanatory Visual Question Answering
1 benchmark · 4 papers with code
Object Hallucination
0 benchmarks · 42 papers with code
Vietnamese Visual Question Answering
0 benchmarks · 4 papers with code
MM-Vet v2
0 benchmarks · 1 paper with code
5 shown of 5 sub-tasks (1 filed under another area).
Few-Shot Learning
27 benchmarks · 1,297 papers with codeFew-Shot Semantic Segmentation
13 benchmarks · 102 papers with code
Few-Shot Audio Classification
10 benchmarks · 5 papers with code
Cross-Domain Few-Shot
9 benchmarks · 80 papers with code
Few-Shot Relation Classification
4 benchmarks · 10 papers with code
One-Shot Learning
1 benchmark · 107 papers with code
5 shown of 12 sub-tasks (3 filed under another area). All sub-tasks of Few-Shot Learning →
Code Generation
27 benchmarks · 745 papers with codeCode Documentation Generation
7 benchmarks · 7 papers with code
Code Translation
2 benchmarks · 54 papers with code
Class-level Code Generation
1 benchmark · 4 papers with code
GitHub issue resolution
0 benchmarks · 6 papers with code
Library-Oriented Code Generation
0 benchmarks · 4 papers with code
5 shown of 5 sub-tasks (1 filed under another area).
Data-to-Text Generation
26 benchmarks · 112 papers with codeKG-to-Text Generation
11 benchmarks · 18 papers with code
Unsupervised KG-to-Text Generation
4 benchmarks · 2 papers with code
Visual Storytelling
1 benchmark · 37 papers with code
3 shown of 3 sub-tasks.
Common Sense Reasoning
24 benchmarks · 325 papers with codeRiddle Sense
2 benchmarks · 5 papers with code
Multiview Contextual Commonsense Inference
2 benchmarks · 1 paper with code
Physical Commonsense Reasoning
1 benchmark · 6 papers with code
Discourse Marker Prediction
1 benchmark · 3 papers with code
Empirical Judgments
1 benchmark · 3 papers with code
5 shown of 17 sub-tasks (1 filed under another area). All sub-tasks of Common Sense Reasoning →
Stance Detection
22 benchmarks · 127 papers with codeStance Detection (US Election 2020 - Biden)
1 benchmark · 1 paper with code
Stance Detection (US Election 2020 - Trump)
1 benchmark · 1 paper with code
Zero-Shot Stance Detection
0 benchmarks · 11 papers with code
Few-Shot Stance Detection
0 benchmarks · 4 papers with code
4 shown of 4 sub-tasks.
Document Classification
21 benchmarks · 235 papers with codePage Stream Segmentation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Text Generation
20 benchmarks · 2,047 papers with codeData-to-Text Generation
26 benchmarks · 112 papers with code
Dialogue Generation
13 benchmarks · 265 papers with code
Table-to-Text Generation
8 benchmarks · 43 papers with code
Code Documentation Generation
7 benchmarks · 7 papers with code
Multi-Document Summarization
5 benchmarks · 113 papers with code
5 shown of 25 sub-tasks (1 filed under another area). All sub-tasks of Text Generation →
Semantic Parsing
20 benchmarks · 414 papers with codeText-To-SQL
10 benchmarks · 209 papers with code
AMR Parsing
8 benchmarks · 51 papers with code
Semantic Dependency Parsing
3 benchmarks · 14 papers with code
DRS Parsing
2 benchmarks · 6 papers with code
UCCA Parsing
2 benchmarks · 4 papers with code
5 shown of 7 sub-tasks. All sub-tasks of Semantic Parsing →
Intent Detection
19 benchmarks · 121 papers with codeOpen Intent Detection
17 benchmarks · 6 papers with code
1 shown of 1 sub-task.
Aspect-Based Sentiment Analysis (ABSA)
18 benchmarks · 185 papers with codeAspect Extraction
6 benchmarks · 34 papers with code
Aspect-Category-Opinion-Sentiment Quadruple Extraction
2 benchmarks · 4 papers with code
Aspect Category Sentiment Analysis
1 benchmark · 12 papers with code
Extract Aspect
1 benchmark · 10 papers with code
Aspect-oriented Opinion Extraction
1 benchmark · 4 papers with code
5 shown of 8 sub-tasks. All sub-tasks of Aspect-Based Sentiment Analysis (ABSA) →
Text-to-Image Generation
17 benchmarks · 546 papers with codeConditional Text-to-Image Synthesis
3 benchmarks · 8 papers with code
Text-based Image Editing
1 benchmark · 29 papers with code
text-guided-image-editing
0 benchmarks · 34 papers with code
Concept Alignment
0 benchmarks · 19 papers with code
Zero-Shot Text-to-Image Generation
0 benchmarks · 11 papers with code
5 shown of 7 sub-tasks. All sub-tasks of Text-to-Image Generation →
Video Generation
16 benchmarks · 609 papers with codeUnconditional Video Generation
1 benchmark · 10 papers with code
Image to Video Generation
0 benchmarks · 38 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Prompt Engineering
16 benchmarks · 454 papers with codeVisual Prompting
0 benchmarks · 60 papers with code
1 shown of 1 sub-task.
Dependency Parsing
16 benchmarks · 338 papers with codeDependency Grammar Induction
2 benchmarks · 3 papers with code
Unsupervised Dependency Parsing
1 benchmark · 5 papers with code
Cross-lingual zero-shot dependency parsing
1 benchmark · 3 papers with code
Transition-Based Dependency Parsing
0 benchmarks · 12 papers with code
Prepositional Phrase Attachment
0 benchmarks · 5 papers with code
5 shown of 5 sub-tasks.
Coreference Resolution
16 benchmarks · 288 papers with codecoreference-resolution
0 benchmarks · 215 papers with code
Cross Document Coreference Resolution
0 benchmarks · 10 papers with code
2 shown of 2 sub-tasks.
Word Sense Disambiguation
16 benchmarks · 150 papers with codeWord Sense Induction
1 benchmark · 19 papers with code
1 shown of 1 sub-task.
Abstractive Text Summarization
15 benchmarks · 362 papers with codeTimeline Summarization
1 benchmark · 8 papers with code
Multimodal Abstractive Text Summarization
1 benchmark · 2 papers with code
Reader-Aware Summarization
1 benchmark · 0 papers with code
3 shown of 3 sub-tasks.
Part-Of-Speech Tagging
15 benchmarks · 228 papers with codeUnsupervised Part-Of-Speech Tagging
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Hate Speech Detection
15 benchmarks · 203 papers with codeHope Speech Detection
2 benchmarks · 9 papers with code
Hate Speech Normalization
0 benchmarks · 3 papers with code
Hate Speech Detection CrisisHateMM Benchmark
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Slot Filling
14 benchmarks · 140 papers with codeZero-shot Slot Filling
3 benchmarks · 10 papers with code
Extracting COVID-19 Events from Twitter
1 benchmark · 2 papers with code
2 shown of 2 sub-tasks.
Semantic Textual Similarity
13 benchmarks · 693 papers with codeParaphrase Identification
11 benchmarks · 76 papers with code
Cross-Lingual Semantic Textual Similarity
0 benchmarks · 3 papers with code
2 shown of 2 sub-tasks.
Dialogue Generation
13 benchmarks · 265 papers with codeMulti-modal Dialogue Generation
1 benchmark · 3 papers with code
1 shown of 1 sub-task.
Cross-Modal Retrieval
13 benchmarks · 244 papers with codeCross-modal retrieval with noisy correspondence
3 benchmarks · 14 papers with code
Image-text matching
1 benchmark · 102 papers with code
Zero-shot Composed Person Retrieval
1 benchmark · 1 paper with code
multilingual cross-modal retrieval
0 benchmarks · 2 papers with code
Cross-Modal Retrieval on RSITMD
0 benchmarks · 0 papers with code
5 shown of 5 sub-tasks.
Grammatical Error Correction
13 benchmarks · 142 papers with codeGrammatical Error Detection
4 benchmarks · 19 papers with code
1 shown of 1 sub-task.
Open Information Extraction
13 benchmarks · 67 papers with codeEvent Extraction
9 benchmarks · 149 papers with code
1 shown of 1 sub-task.
Retrieval
11 benchmarks · 5,274 papers with codeText Retrieval
16 benchmarks · 335 papers with code
Table Retrieval
1 benchmark · 15 papers with code
Deep Hashing
0 benchmarks · 56 papers with code
3 shown of 3 sub-tasks.
Information Retrieval
11 benchmarks · 1,188 papers with codePassage Retrieval
5 benchmarks · 133 papers with code
Scientific Results Extraction
2 benchmarks · 5 papers with code
Zero Shot on BEIR (Inference Free Model)
1 benchmark · 4 papers with code
TAR
0 benchmarks · 35 papers with code
Cross-Lingual Information Retrieval
0 benchmarks · 14 papers with code
5 shown of 6 sub-tasks. All sub-tasks of Information Retrieval →
Question Generation
11 benchmarks · 264 papers with codePoll Generation
1 benchmark · 3 papers with code
1 shown of 1 sub-task.
Speech-to-Text Translation
11 benchmarks · 64 papers with codeSimultaneous Speech-to-Text Translation
0 benchmarks · 3 papers with code
1 shown of 1 sub-task.
Entity Resolution
11 benchmarks · 55 papers with codeBlocking
5 benchmarks · 129 papers with code
1 shown of 1 sub-task.
2D Semantic Segmentation
11 benchmarks · 49 papers with codeImage Segmentation
13 benchmarks · 2,073 papers with code
Human Part Segmentation
6 benchmarks · 15 papers with code
Reflection Removal
5 benchmarks · 38 papers with code
Continual Semantic Segmentation
3 benchmarks · 17 papers with code
Text Style Transfer
2 benchmarks · 93 papers with code
5 shown of 17 sub-tasks (5 filed under another area). All sub-tasks of 2D Semantic Segmentation →
Mathematical Reasoning
10 benchmarks · 395 papers with codeMath Word Problem Solving
13 benchmarks · 80 papers with code
Formal Logic
1 benchmark · 20 papers with code
Abstract Algebra
1 benchmark · 6 papers with code
Mathematical Induction
1 benchmark · 4 papers with code
High School Mathematics
1 benchmark · 1 paper with code
5 shown of 7 sub-tasks (2 filed under another area). All sub-tasks of Mathematical Reasoning →
Text-To-SQL
10 benchmarks · 209 papers with codeMMSQL performance
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Entity Alignment
10 benchmarks · 117 papers with codeMulti-modal Entity Alignment
8 benchmarks · 14 papers with code
1 shown of 1 sub-task (1 filed under another area).
Cross-Lingual Document Classification
10 benchmarks · 12 papers with codeNews Classification
4 benchmarks · 33 papers with code
1 shown of 1 sub-task.
Emotion Recognition
9 benchmarks · 614 papers with codeSpeech Emotion Recognition
16 benchmarks · 139 papers with code
Emotion Recognition in Conversation
16 benchmarks · 83 papers with code
Multimodal Emotion Recognition
7 benchmarks · 80 papers with code
Emotion Recognition in Context
4 benchmarks · 5 papers with code
EEG Emotion Recognition
3 benchmarks · 14 papers with code
5 shown of 12 sub-tasks (2 filed under another area). All sub-tasks of Emotion Recognition →
Event Extraction
9 benchmarks · 149 papers with codeNER
4 benchmarks · 665 papers with code
Event Causality Identification
0 benchmarks · 13 papers with code
Zero-shot Event Extraction
0 benchmarks · 4 papers with code
3 shown of 3 sub-tasks.
Source Code Summarization
9 benchmarks · 39 papers with codeMethod name prediction
1 benchmark · 15 papers with code
1 shown of 1 sub-task.
Relation Classification
8 benchmarks · 160 papers with codeFew-Shot Relation Classification
4 benchmarks · 10 papers with code
Implicit Discourse Relation Classification
0 benchmarks · 7 papers with code
Cause-Effect Relation Classification
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Entity Typing
8 benchmarks · 92 papers with codeEntity Typing on DH-KGs
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Few-Shot Text Classification
8 benchmarks · 46 papers with codeZero-Shot Out-of-Domain Detection
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Table-to-Text Generation
8 benchmarks · 43 papers with codeKB-to-Language Generation
1 benchmark · 3 papers with code
1 shown of 1 sub-task.
Knowledge Distillation
7 benchmarks · 1,740 papers with codeData-free Knowledge Distillation
2 benchmarks · 37 papers with code
Self-Knowledge Distillation
0 benchmarks · 36 papers with code
2 shown of 2 sub-tasks.
Reading Comprehension
7 benchmarks · 634 papers with codeMachine Reading Comprehension
4 benchmarks · 206 papers with code
Intent Recognition
1 benchmark · 22 papers with code
Implicit Relations
1 benchmark · 20 papers with code
Question Selection
1 benchmark · 18 papers with code
LAMBADA
1 benchmark · 15 papers with code
5 shown of 20 sub-tasks (1 filed under another area). All sub-tasks of Reading Comprehension →
Semantic Similarity
7 benchmarks · 534 papers with codeSemantic Shift Detection
0 benchmarks · 5 papers with code
Similarity Explanation
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Knowledge Graph Completion
7 benchmarks · 251 papers with codeInductive knowledge graph completion
3 benchmarks · 14 papers with code
Triple Classification
1 benchmark · 23 papers with code
Large Language Model
0 benchmarks · 2,250 papers with code
Inductive Relation Prediction
0 benchmarks · 11 papers with code
4 shown of 4 sub-tasks (1 filed under another area).
Document Summarization
7 benchmarks · 226 papers with codeEmail Thread Summarization
2 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Semantic Role Labeling
7 benchmarks · 138 papers with codePredicate Detection
3 benchmarks · 9 papers with code
Semantic Role Labeling (predicted predicates)
2 benchmarks · 4 papers with code
Textual Analogy Parsing
0 benchmarks · 2 papers with code
3 shown of 3 sub-tasks.
Binary text classification
7 benchmarks · 12 papers with codeDetection of potentially void clauses
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Natural Language Understanding
6 benchmarks · 809 papers with codeVietnamese Social Media Text Processing
0 benchmarks · 9 papers with code
Emotional Dialogue Acts
0 benchmarks · 2 papers with code
2 shown of 2 sub-tasks.
Optical Character Recognition (OCR)
6 benchmarks · 462 papers with codeHandwritten Text Recognition
13 benchmarks · 56 papers with code
Handwriting Recognition
3 benchmarks · 53 papers with code
Handwritten Digit Recognition
2 benchmarks · 29 papers with code
Active Learning
1 benchmark · 913 papers with code
Irregular Text Recognition
0 benchmarks · 5 papers with code
5 shown of 10 sub-tasks (2 filed under another area). All sub-tasks of Optical Character Recognition (OCR) →
Topic Models
6 benchmarks · 229 papers with codeTopic coverage
3 benchmarks · 6 papers with code
Dynamic Topic Modeling
0 benchmarks · 3 papers with code
2 shown of 2 sub-tasks.
Intrusion Detection
6 benchmarks · 151 papers with codeNetwork Intrusion Detection
6 benchmarks · 67 papers with code
1 shown of 1 sub-task.
Language Identification
6 benchmarks · 143 papers with codeNative Language Identification
1 benchmark · 5 papers with code
Dialect Identification
0 benchmarks · 33 papers with code
2 shown of 2 sub-tasks.
Sentence Classification
6 benchmarks · 115 papers with codeUnfairness Detection
0 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Text-To-Speech Synthesis
6 benchmarks · 104 papers with codeProsody Prediction
1 benchmark · 4 papers with code
Zero-Shot Multi-Speaker TTS
0 benchmarks · 3 papers with code
2 shown of 2 sub-tasks.
Text-to-Video Generation
6 benchmarks · 97 papers with codeText-to-Video Editing
0 benchmarks · 6 papers with code
Subject-driven Video Generation
0 benchmarks · 2 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Key Information Extraction
6 benchmarks · 40 papers with codeKey-value Pair Extraction
2 benchmarks · 9 papers with code
1 shown of 1 sub-task.
Aspect Extraction
6 benchmarks · 34 papers with codeHidden Aspect Detection
0 benchmarks · 2 papers with code
Latent Aspect Detection
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Representation Learning
5 benchmarks · 4,662 papers with codeDisentanglement
3 benchmarks · 744 papers with code
Graph Representation Learning
1 benchmark · 479 papers with code
Feature Upsampling
1 benchmark · 22 papers with code
Sentence Embeddings
0 benchmarks · 255 papers with code
Network Embedding
0 benchmarks · 165 papers with code
5 shown of 14 sub-tasks (2 filed under another area). All sub-tasks of Representation Learning →
Deep Clustering
5 benchmarks · 135 papers with codeTrajectory Clustering
0 benchmarks · 6 papers with code
Deep Nonparametric Clustering
0 benchmarks · 1 paper with code
NONPARAMETRIC DEEP CLUSTERING
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Story Generation
5 benchmarks · 110 papers with codeVisual Storytelling
1 benchmark · 37 papers with code
1 shown of 1 sub-task.
Bias Detection
5 benchmarks · 80 papers with codeSelection bias
0 benchmarks · 143 papers with code
1 shown of 1 sub-task.
Phrase Grounding
5 benchmarks · 51 papers with codeGrounded Open Vocabulary Acquisition
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Extractive Text Summarization
5 benchmarks · 34 papers with codeReader-Aware Summarization
1 benchmark · 0 papers with code
1 shown of 1 sub-task.
Binary Classification
4 benchmarks · 710 papers with codeCancer-no cancer per breast classification
2 benchmarks · 2 papers with code
Cancer-no cancer per image classification
1 benchmark · 2 papers with code
Stable MCI vs Progressive MCI
1 benchmark · 1 paper with code
Suspicous (BIRADS 4,5)-no suspicous (BIRADS 1,2,3) per image classification
1 benchmark · 1 paper with code
LLM-generated Text Detection
0 benchmarks · 11 papers with code
5 shown of 6 sub-tasks. All sub-tasks of Binary Classification →
Task-Oriented Dialogue Systems
4 benchmarks · 131 papers with codeSSTOD
2 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Constituency Parsing
4 benchmarks · 81 papers with codeConstituency Grammar Induction
4 benchmarks · 20 papers with code
1 shown of 1 sub-task.
Document-level Relation Extraction
4 benchmarks · 66 papers with codeDocument-level RE with incomplete labeling
2 benchmarks · 3 papers with code
1 shown of 1 sub-task.
Discourse Parsing
4 benchmarks · 57 papers with codeEnd-to-End RST Parsing
1 benchmark · 3 papers with code
Discourse Segmentation
0 benchmarks · 14 papers with code
Connective Detection
0 benchmarks · 2 papers with code
3 shown of 3 sub-tasks.
Document Text Classification
4 benchmarks · 5 papers with codeLearning with noisy labels
20 benchmarks · 143 papers with code
Multi-Label Classification Of Biomedical Texts
2 benchmarks · 4 papers with code
Political Salient Issue Orientation Detection
1 benchmark · 3 papers with code
3 shown of 3 sub-tasks (1 filed under another area).
Data Augmentation
3 benchmarks · 3,225 papers with codeImage Augmentation
1 benchmark · 127 papers with code
Text Augmentation
0 benchmarks · 41 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Paraphrase Generation
3 benchmarks · 74 papers with codeMultilingual Paraphrase Generation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
3D Action Recognition
3 benchmarks · 38 papers with codeSkeleton Based Action Recognition
34 benchmarks · 219 papers with code
Image Manipulation Detection
21 benchmarks · 38 papers with code
Zero Shot Skeletal Action Recognition
3 benchmarks · 7 papers with code
Generalized Zero Shot skeletal action recognition
3 benchmarks · 3 papers with code
Model Editing
0 benchmarks · 107 papers with code
5 shown of 6 sub-tasks (2 filed under another area). All sub-tasks of 3D Action Recognition →
Multimodal Machine Translation
3 benchmarks · 38 papers with codeMultimodal Lexical Translation
4 benchmarks · 2 papers with code
Face to Face Translation
0 benchmarks · 4 papers with code
2 shown of 2 sub-tasks.
Text Clustering
3 benchmarks · 38 papers with codeShort Text Clustering
8 benchmarks · 18 papers with code
Open Intent Discovery
6 benchmarks · 5 papers with code
Hierarchical Text Clustering
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Meme Classification
3 benchmarks · 30 papers with codeHateful Meme Classification
4 benchmarks · 9 papers with code
1 shown of 1 sub-task.
Age And Gender Classification
3 benchmarks · 16 papers with codeAuthor Profiling
0 benchmarks · 14 papers with code
1 shown of 1 sub-task.
Inductive knowledge graph completion
3 benchmarks · 14 papers with codeLarge Language Model
0 benchmarks · 2,250 papers with code
1 shown of 1 sub-task.
Biomedical Information Retrieval
3 benchmarks · 7 papers with codeSpO2 estimation
3 benchmarks · 3 papers with code
PICO
1 benchmark · 26 papers with code
2 shown of 2 sub-tasks.
Multimodal Text and Image Classification
3 benchmarks · 5 papers with codeimage-sentence alignment
12 benchmarks · 2 papers with code
Open-World Social Event Classification
0 benchmarks · 0 papers with code
2 shown of 2 sub-tasks.
Text Style Transfer
2 benchmarks · 93 papers with codeFormality Style Transfer
1 benchmark · 10 papers with code
Semi-Supervised Formality Style Transfer
0 benchmarks · 1 paper with code
Word Attribute Transfer
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Document Ranking
2 benchmarks · 65 papers with codeSession Search
0 benchmarks · 2 papers with code
1 shown of 1 sub-task.
Term Extraction
2 benchmarks · 44 papers with codeNested Term Extraction
3 benchmarks · 2 papers with code
Nested Term Recognition
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Data-free Knowledge Distillation
2 benchmarks · 37 papers with codeBenchmarking
2 benchmarks · 2,658 papers with code
1 shown of 1 sub-task (1 filed under another area).
Weakly Supervised Classification
2 benchmarks · 25 papers with codeWeakly Supervised Data Denoising
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Temporal Information Extraction
2 benchmarks · 20 papers with codeTemporal Tagging
8 benchmarks · 4 papers with code
1 shown of 1 sub-task.
Hope Speech Detection
2 benchmarks · 9 papers with codeHope Speech Detection for English
1 benchmark · 1 paper with code
Hope Speech Detection for Malayalam
1 benchmark · 1 paper with code
Hope Speech Detection for Tamil
1 benchmark · 1 paper with code
3 shown of 3 sub-tasks.
Handwriting Verification
2 benchmarks · 6 papers with codeBangla Spelling Error Correction
1 benchmark · 3 papers with code
1 shown of 1 sub-task.
Aspect-Category-Opinion-Sentiment Quadruple Extraction
2 benchmarks · 4 papers with codeConversational Sentiment Quadruple Extraction
2 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Recognizing Emotion Cause in Conversations
2 benchmarks · 2 papers with codeCausal Emotion Entailment
1 benchmark · 10 papers with code
1 shown of 1 sub-task.
Reinforcement Learning
1 benchmark · 4,183 papers with codeDeep Reinforcement Learning
0 benchmarks · 1,739 papers with code
1 shown of 1 sub-task.
Active Learning
1 benchmark · 913 papers with codeActive Object Detection
2 benchmarks · 6 papers with code
1 shown of 1 sub-task.
Instruction Following
1 benchmark · 609 papers with codevisual instruction following
1 benchmark · 14 papers with code
1 shown of 1 sub-task.
Cross-Lingual Transfer
1 benchmark · 333 papers with codeCross-Lingual NER
28 benchmarks · 27 papers with code
Zero-Shot Cross-Lingual Transfer
2 benchmarks · 80 papers with code
2 shown of 2 sub-tasks.
Chatbot
1 benchmark · 269 papers with codeDialogue Generation
13 benchmarks · 265 papers with code
1 shown of 1 sub-task.
Natural Language Queries
1 benchmark · 135 papers with codeText-to-CQL
0 benchmarks · 0 papers with code
1 shown of 1 sub-task.
Argument Mining
1 benchmark · 98 papers with codeValNov
2 benchmarks · 1 paper with code
Component Classification
1 benchmark · 8 papers with code
Argument Pair Extraction (APE)
1 benchmark · 4 papers with code
Claim Extraction with Stance Classification (CESC)
1 benchmark · 1 paper with code
Claim-Evidence Pair Extraction (CEPE)
1 benchmark · 1 paper with code
5 shown of 6 sub-tasks. All sub-tasks of Argument Mining →
Multimodal Deep Learning
1 benchmark · 97 papers with codeMultimodal Text and Image Classification
3 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Language Acquisition
1 benchmark · 83 papers with codeGrounded language learning
0 benchmarks · 23 papers with code
1 shown of 1 sub-task.
Conversational Question Answering
1 benchmark · 63 papers with codeQuestion Rewriting
0 benchmarks · 15 papers with code
1 shown of 1 sub-task.
Sentence Completion
1 benchmark · 49 papers with codeHurtful Sentence Completion
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Spam detection
1 benchmark · 38 papers with codeTraditional Spam Detection
1 benchmark · 2 papers with code
Context-specific Spam Detection
1 benchmark · 1 paper with code
2 shown of 2 sub-tasks.
Visual Storytelling
1 benchmark · 37 papers with codeImage-guided Story Ending Generation
2 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Conditional Text Generation
1 benchmark · 32 papers with codeContextualized Literature-based Discovery
0 benchmarks · 1 paper with code
Multimedia Generative Script Learning
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Open-Domain Dialog
1 benchmark · 32 papers with codeDialogue Evaluation
2 benchmarks · 61 papers with code
1 shown of 1 sub-task.
Temporal Relation Extraction
1 benchmark · 32 papers with codeTemporal Relation Classification
4 benchmarks · 7 papers with code
1 shown of 1 sub-task.
Ad-Hoc Information Retrieval
1 benchmark · 27 papers with codeDocument Ranking
2 benchmarks · 65 papers with code
1 shown of 1 sub-task.
Document AI
1 benchmark · 24 papers with codedocument understanding
0 benchmarks · 140 papers with code
1 shown of 1 sub-task.
Goal-Oriented Dialog
1 benchmark · 24 papers with codeUser Simulation
0 benchmarks · 25 papers with code
1 shown of 1 sub-task.
Sentence Compression
1 benchmark · 22 papers with codeUnsupervised Sentence Compression
0 benchmarks · 4 papers with code
1 shown of 1 sub-task.
Emotional Intelligence
1 benchmark · 19 papers with codeDark Humor Detection
1 benchmark · 2 papers with code
SNARKS
0 benchmarks · 6 papers with code
Ruin Names
0 benchmarks · 3 papers with code
3 shown of 3 sub-tasks.
Question Similarity
1 benchmark · 17 papers with codeMedical question pair similarity computation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Lexical Normalization
1 benchmark · 16 papers with codePronunciation Dictionary Creation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Attribute Extraction
1 benchmark · 11 papers with codelegal outcome extraction
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Clinical Concept Extraction
1 benchmark · 9 papers with codeClinical Information Retreival
0 benchmarks · 2 papers with code
1 shown of 1 sub-task.
Scientific Document Summarization
1 benchmark · 9 papers with codeLay Summarization
2 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Summarization
1 benchmark · 4 papers with codeUnsupervised Extractive Summarization
3 benchmarks · 18 papers with code
Query-focused Summarization
0 benchmarks · 18 papers with code
2 shown of 2 sub-tasks.
1 Image, 2*2 Stitchi
1 benchmark · 3 papers with codePose Estimation
31 benchmarks · 1,679 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Image Deblurring
9 benchmarks · 167 papers with code
Virtual Try-on
9 benchmarks · 114 papers with code
Style Transfer
3 benchmarks · 759 papers with code
5 shown of 12 sub-tasks (9 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →
Information Extraction
1 benchmark · 3 papers with codeJoint Entity and Relation Extraction
16 benchmarks · 56 papers with code
Event Extraction
9 benchmarks · 149 papers with code
Attribute Value Extraction
4 benchmarks · 16 papers with code
Low Resource Named Entity Recognition
3 benchmarks · 15 papers with code
Drug–drug Interaction Extraction
3 benchmarks · 13 papers with code
5 shown of 16 sub-tasks (1 filed under another area). All sub-tasks of Information Extraction →
Text2text Generation
1 benchmark · 2 papers with codeKeyphrase Generation
0 benchmarks · 43 papers with code
Figurative Language Visualization
0 benchmarks · 1 paper with code
Sketch-to-text Generation
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Dialogue
1 benchmark · 1 paper with codeDialogue Generation
13 benchmarks · 265 papers with code
Visual Dialog
8 benchmarks · 56 papers with code
Dialogue State Tracking
7 benchmarks · 138 papers with code
Dialogue Act Classification
5 benchmarks · 23 papers with code
Task-Oriented Dialogue Systems
4 benchmarks · 131 papers with code
5 shown of 21 sub-tasks (1 filed under another area). All sub-tasks of Dialogue →
Language Modeling
0 benchmarks · 5,620 papers with codeDream Generation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Deep Learning
0 benchmarks · 2,693 papers with codePolynomial Neural Networks
0 benchmarks · 10 papers with code
1 shown of 1 sub-task.
Large Language Model
0 benchmarks · 2,250 papers with codeKnowledge Graphs
4 benchmarks · 1,273 papers with code
RAG
1 benchmark · 758 papers with code
AI Agent
0 benchmarks · 111 papers with code
3 shown of 3 sub-tasks (2 filed under another area).
Word Embeddings
0 benchmarks · 1,177 papers with codeLearning Word Embeddings
0 benchmarks · 24 papers with code
Multilingual Word Embeddings
0 benchmarks · 19 papers with code
Embeddings Evaluation
0 benchmarks · 10 papers with code
Contextualised Word Representations
0 benchmarks · 4 papers with code
4 shown of 4 sub-tasks.
NMT
0 benchmarks · 523 papers with codeDirect NMT
0 benchmarks · 0 papers with code
1 shown of 1 sub-task.
Text to Image Generation
0 benchmarks · 461 papers with codeText to 3D
1 benchmark · 102 papers with code
1 shown of 1 sub-task (1 filed under another area).
Sentence Embeddings
0 benchmarks · 255 papers with codeSentence Embeddings For Biomedical Texts
2 benchmarks · 2 papers with code
Sentence Compression
1 benchmark · 22 papers with code
Sentence Embedding
0 benchmarks · 162 papers with code
Joint Multilingual Sentence Representations
0 benchmarks · 2 papers with code
4 shown of 4 sub-tasks.
Symbolic Regression
0 benchmarks · 155 papers with codeEquation Discovery
0 benchmarks · 34 papers with code
1 shown of 1 sub-task.
document understanding
0 benchmarks · 140 papers with codeLine Items Extraction
0 benchmarks · 0 papers with code
1 shown of 1 sub-task.
Data Integration
0 benchmarks · 130 papers with codeEntity Resolution
11 benchmarks · 55 papers with code
Entity Alignment
10 benchmarks · 117 papers with code
Table annotation
0 benchmarks · 23 papers with code
3 shown of 3 sub-tasks.
Model Editing
0 benchmarks · 107 papers with codeknowledge editing
1 benchmark · 88 papers with code
1 shown of 1 sub-task.
Authorship Attribution
0 benchmarks · 62 papers with codeSource Code Authorship Attribution
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
De-identification
0 benchmarks · 54 papers with codePrivacy Preserving Deep Learning
0 benchmarks · 30 papers with code
Full-body anonymization
0 benchmarks · 4 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Spelling Correction
0 benchmarks · 52 papers with codeBangla Spelling Error Correction
1 benchmark · 3 papers with code
1 shown of 1 sub-task.
Token Classification
0 benchmarks · 47 papers with codeBlackout Poetry Generation
1 benchmark · 1 paper with code
Toxic Spans Detection
0 benchmarks · 18 papers with code
2 shown of 2 sub-tasks.
Dialogue Understanding
0 benchmarks · 35 papers with codeSpoken Language Understanding
5 benchmarks · 135 papers with code
Dialogue Safety Prediction
2 benchmarks · 3 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Abuse Detection
0 benchmarks · 32 papers with codeHate Speech Detection
15 benchmarks · 203 papers with code
1 shown of 1 sub-task.
Constrained Clustering
0 benchmarks · 29 papers with codeIncremental Constrained Clustering
2 benchmarks · 1 paper with code
Only Connect Walls Dataset Task 1 (Grouping)
1 benchmark · 10 papers with code
2 shown of 2 sub-tasks.
Stock Prediction
0 benchmarks · 29 papers with codePAIR TRADING
2 benchmarks · 2 papers with code
Text-Based Stock Prediction
0 benchmarks · 4 papers with code
Event-Driven Trading
0 benchmarks · 1 paper with code
3 shown of 3 sub-tasks.
Table annotation
0 benchmarks · 23 papers with codeColumn Type Annotation
12 benchmarks · 19 papers with code
Cell Entity Annotation
5 benchmarks · 6 papers with code
Columns Property Annotation
4 benchmarks · 5 papers with code
Row Annotation
1 benchmark · 1 paper with code
Table Type Detection
1 benchmark · 1 paper with code
5 shown of 6 sub-tasks. All sub-tasks of Table annotation →
Sentence Summarization
0 benchmarks · 19 papers with codeUnsupervised Sentence Summarization
0 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Conversational Response Generation
0 benchmarks · 17 papers with codePersonalized and Emotional Conversation
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Propaganda detection
0 benchmarks · 15 papers with codePropaganda span identification
0 benchmarks · 1 paper with code
Propaganda technique identification
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Twitter Sentiment Analysis
0 benchmarks · 14 papers with codeTweet-Reply Sentiment Analysis
1 benchmark · 0 papers with code
1 shown of 1 sub-task.
News Generation
0 benchmarks · 12 papers with codeHeadline Generation
1 benchmark · 36 papers with code
1 shown of 1 sub-task.
Negation Detection
0 benchmarks · 11 papers with codeNegation Scope Resolution
4 benchmarks · 4 papers with code
1 shown of 1 sub-task.
Cross-Lingual Entity Linking
0 benchmarks · 9 papers with codeVariable Disambiguation
1 benchmark · 1 paper with code
1 shown of 1 sub-task.
Lexical Analysis
0 benchmarks · 8 papers with codeLexical Complexity Prediction
0 benchmarks · 11 papers with code
1 shown of 1 sub-task.
Literature Mining
0 benchmarks · 7 papers with codeSystematic Literature Review
0 benchmarks · 49 papers with code
1 shown of 1 sub-task.
Persian Sentiment Analysis
0 benchmarks · 7 papers with codeTransition-Based Dependency Parsing
0 benchmarks · 12 papers with code
1 shown of 1 sub-task.
Sentence Pair Modeling
0 benchmarks · 6 papers with codeSemantic Similarity
7 benchmarks · 534 papers with code
1 shown of 1 sub-task.
Data Mining
0 benchmarks · 4 papers with codeArgument Mining
1 benchmark · 98 papers with code
Opinion Mining
1 benchmark · 64 papers with code
Sequential Pattern Mining
1 benchmark · 9 papers with code
CSV dialect detection
1 benchmark · 1 paper with code
cognitive diagnosis
0 benchmarks · 22 papers with code
5 shown of 7 sub-tasks. All sub-tasks of Data Mining →
Continual Named Entity Recognition
0 benchmarks · 3 papers with codeFG-1-PG-1
3 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Multimodal Association
0 benchmarks · 3 papers with codemultimodal generation
1 benchmark · 53 papers with code
1 shown of 1 sub-task.
Speculation Detection
0 benchmarks · 3 papers with codeSpeculation Scope Resolution
3 benchmarks · 2 papers with code
1 shown of 1 sub-task.
Joint Multilingual Sentence Representations
0 benchmarks · 2 papers with codeAbstract Meaning Representation
0 benchmarks · 92 papers with code
1 shown of 1 sub-task.
Natural Language Transduction
0 benchmarks · 2 papers with codeLipreading
8 benchmarks · 36 papers with code
1 shown of 1 sub-task (1 filed under another area).
Personality Generation
0 benchmarks · 2 papers with codePersonality Alignment
0 benchmarks · 3 papers with code
1 shown of 1 sub-task.
Political evalutation
0 benchmarks · 2 papers with codeAlignement visualisation
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
trustable and focussed LLM generated content
0 benchmarks · 2 papers with codeGame Design
0 benchmarks · 14 papers with code
1 shown of 1 sub-task.
Vietnamese Aspect-Based Sentiment Analysis
0 benchmarks · 2 papers with codeSentiment Dependency Learning
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Vietnamese Sentiment Analysis
0 benchmarks · 2 papers with codeVietnamese Multimodal Sentiment Analysis
0 benchmarks · 1 paper with code
1 shown of 1 sub-task.
Nested Term Recognition
0 benchmarks · 1 paper with codeNested Term Recognition from Flat Supervision
3 benchmarks · 1 paper with code
1 shown of 1 sub-task.
NLP based Person Retrival
0 benchmarks · 1 paper with codeDecoder
1 benchmark · 4,358 papers with code
1 shown of 1 sub-task.
Anaphora Resolution
0 benchmarks · 0 papers with codeAbstract Anaphora Resolution
1 benchmark · 1 paper with code
Bridging Anaphora Resolution
0 benchmarks · 2 papers with code
2 shown of 2 sub-tasks.
Chinese
0 benchmarks · 0 papers with codeChinese Word Segmentation
6 benchmarks · 50 papers with code
Handwritten Chinese Text Recognition
0 benchmarks · 3 papers with code
Chinese Spelling Error Correction
0 benchmarks · 2 papers with code
Chinese Zero Pronoun Resolution
0 benchmarks · 2 papers with code
Offline Handwritten Chinese Character Recognition
0 benchmarks · 2 papers with code
5 shown of 5 sub-tasks.
Cross-Lingual
0 benchmarks · 0 papers with codeCross-Lingual Document Classification
10 benchmarks · 12 papers with code
Cross-Lingual Transfer
1 benchmark · 333 papers with code
Cross-Lingual Entity Linking
0 benchmarks · 9 papers with code
Cross-Language Text Summarization
0 benchmarks · 0 papers with code
4 shown of 4 sub-tasks.
Optical Charater Recogntion
0 benchmarks · 0 papers with codeBangla Text Detection
1 benchmark · 0 papers with code
1 shown of 1 sub-task.
Pcl Detection
0 benchmarks · 0 papers with codeSemEval-2022 Task 4-1 (Binary PCL Detection)
1 benchmark · 0 papers with code
SemEval-2022 Task 4-2 (Multi-label PCL Detection)
1 benchmark · 0 papers with code
2 shown of 2 sub-tasks.
Shallow Syntax
0 benchmarks · 0 papers with codeChunking
5 benchmarks · 120 papers with code
1 shown of 1 sub-task.
Taxonomy Learning
0 benchmarks · 0 papers with codeHypernym Discovery
3 benchmarks · 8 papers with code
Taxonomy Expansion
0 benchmarks · 13 papers with code
2 shown of 2 sub-tasks.
Temporal Processing
0 benchmarks · 0 papers with codeTemporal Information Extraction
2 benchmarks · 20 papers with code
Timex normalization
2 benchmarks · 5 papers with code
Document Dating
2 benchmarks · 2 papers with code
3 shown of 3 sub-tasks.
Tasks with no parent task
224 tasks in Natural Language Processing sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.
Entity Linking
27 benchmarks · 277 papers with code
Conversational Response Selection
14 benchmarks · 36 papers with code
Text Simplification
11 benchmarks · 129 papers with code
Entity Disambiguation
11 benchmarks · 63 papers with code
Fake News Detection
10 benchmarks · 203 papers with code
Sarcasm Detection
9 benchmarks · 76 papers with code
Image-to-Text Retrieval
8 benchmarks · 37 papers with code
Word Alignment
7 benchmarks · 92 papers with code
Explanation Generation
5 benchmarks · 92 papers with code
Long-Context Understanding
5 benchmarks · 54 papers with code
Keyphrase Extraction
5 benchmarks · 52 papers with code
Linguistic Acceptability
5 benchmarks · 49 papers with code
Recipe Generation
5 benchmarks · 15 papers with code
Intent Classification
4 benchmarks · 113 papers with code
Aspect Category Detection
4 benchmarks · 13 papers with code
Extreme Summarization
4 benchmarks · 13 papers with code
Cross-Lingual Bitext Mining
4 benchmarks · 5 papers with code
Translation
3 benchmarks · 3,574 papers with code
Response Generation
3 benchmarks · 370 papers with code
Fact Verification
3 benchmarks · 129 papers with code
Keyword Extraction
3 benchmarks · 33 papers with code
Intent Discovery
3 benchmarks · 20 papers with code
Dialogue Rewriting
3 benchmarks · 4 papers with code
Attribute Mining
3 benchmarks · 3 papers with code
POS Tagging
2 benchmarks · 136 papers with code
Legal Reasoning
2 benchmarks · 30 papers with code
Graph-to-Sequence
2 benchmarks · 29 papers with code
Cloze Test
2 benchmarks · 28 papers with code
Rumour Detection
2 benchmarks · 23 papers with code
Subjectivity Analysis
2 benchmarks · 21 papers with code
Passage Re-Ranking
2 benchmarks · 19 papers with code
Meeting Summarization
2 benchmarks · 18 papers with code
Semantic entity labeling
2 benchmarks · 12 papers with code
Nested Mention Recognition
2 benchmarks · 10 papers with code
Arabic Text Diacritization
2 benchmarks · 7 papers with code
Thai Word Segmentation
2 benchmarks · 7 papers with code
Reading Order Detection
2 benchmarks · 4 papers with code
Crowdsourced Text Aggregation
2 benchmarks · 1 paper with code
Negation and Speculation Cue Detection
2 benchmarks · 1 paper with code
Phrase Ranking
2 benchmarks · 1 paper with code
Phrase Tagging
2 benchmarks · 1 paper with code
Memorization
1 benchmark · 438 papers with code
GSM8K
1 benchmark · 209 papers with code
Relational Reasoning
1 benchmark · 179 papers with code
Word Similarity
1 benchmark · 117 papers with code
Passage Ranking
1 benchmark · 35 papers with code
Knowledge Base Population
1 benchmark · 32 papers with code
Automated Essay Scoring
1 benchmark · 27 papers with code
Semantic Retrieval
1 benchmark · 26 papers with code
Sentence Ordering
1 benchmark · 21 papers with code
Humor Detection
1 benchmark · 20 papers with code
Table-based Fact Verification
1 benchmark · 19 papers with code
Code Repair
1 benchmark · 14 papers with code
Dialog Act Classification
1 benchmark · 11 papers with code
Probing Language Models
1 benchmark · 11 papers with code
CCG Supertagging
1 benchmark · 8 papers with code
Fact Selection
1 benchmark · 7 papers with code
Multi-agent Integration
1 benchmark · 7 papers with code
answerability prediction
1 benchmark · 6 papers with code
Aspect Category Polarity
1 benchmark · 5 papers with code
Chinese Spell Checking
1 benchmark · 4 papers with code
Stereotypical Bias Analysis
1 benchmark · 4 papers with code
Zero-shot Sentiment Classification
1 benchmark · 4 papers with code
Action Parsing
1 benchmark · 3 papers with code
Domain Labelling
1 benchmark · 3 papers with code
Japanese Word Segmentation
1 benchmark · 3 papers with code
Memex Question Answering
1 benchmark · 3 papers with code
Polyphone disambiguation
1 benchmark · 3 papers with code
Twitter Event Detection
1 benchmark · 3 papers with code
AMR Graph Similarity
1 benchmark · 2 papers with code
Binary Condescension Detection
1 benchmark · 2 papers with code
Conversational Web Navigation
1 benchmark · 2 papers with code
Croatian Text Diacritization
1 benchmark · 2 papers with code
Czech Text Diacritization
1 benchmark · 2 papers with code
Description-guided molecule generation
1 benchmark · 2 papers with code
French Text Diacritization
1 benchmark · 2 papers with code
Hungarian Text Diacritization
1 benchmark · 2 papers with code
Irish Text Diacritization
1 benchmark · 2 papers with code
Latvian Text Diacritization
1 benchmark · 2 papers with code
Morpheme Segmentaiton
1 benchmark · 2 papers with code
Multi-label Condescension Detection
1 benchmark · 2 papers with code
Personality Recognition in Conversation
1 benchmark · 2 papers with code
Role-filler Entity Extraction
1 benchmark · 2 papers with code
Romanian Text Diacritization
1 benchmark · 2 papers with code
Slovak Text Diacritization
1 benchmark · 2 papers with code
Spanish Text Diacritization
1 benchmark · 2 papers with code
Turkish Text Diacritization
1 benchmark · 2 papers with code
Vietnamese Text Diacritization
1 benchmark · 2 papers with code
Chemical Indexing
1 benchmark · 1 paper with code
Clinical Assertion Status Detection
1 benchmark · 1 paper with code
Commonsense Reasoning for RL
1 benchmark · 1 paper with code
GermEval2024 Shared Task 1 Subtask 1
1 benchmark · 1 paper with code
GermEval2024 Shared Task 1 Subtask 2
1 benchmark · 1 paper with code
Math Information Retrieval
1 benchmark · 1 paper with code
Multimodal Text Prediction
1 benchmark · 1 paper with code
Poem meters classification
1 benchmark · 1 paper with code
Query Wellformedness
1 benchmark · 1 paper with code
Question-Answer categorization
1 benchmark · 1 paper with code
Speaker Attribution in German Parliamentary Debates (GermEval 2023, subtask 1)
1 benchmark · 1 paper with code
TinyQA Benchmark++
1 benchmark · 1 paper with code
Counterspeech Detection
1 benchmark · 0 papers with code
In-Context Learning
0 benchmarks · 998 papers with code
Retrieval-augmented Generation
0 benchmarks · 777 papers with code
Mamba
0 benchmarks · 552 papers with code
Specificity
0 benchmarks · 504 papers with code
Drug Design
0 benchmarks · 180 papers with code
Text Matching
0 benchmarks · 161 papers with code
Safety Alignment
0 benchmarks · 134 papers with code
Self-Learning
0 benchmarks · 125 papers with code
molecular representation
0 benchmarks · 87 papers with code
Novelty Detection
0 benchmarks · 84 papers with code
Conversational Search
0 benchmarks · 75 papers with code
Morphological Analysis
0 benchmarks · 73 papers with code
Lemmatization
0 benchmarks · 68 papers with code
Abusive Language
0 benchmarks · 51 papers with code
Protein Folding
0 benchmarks · 51 papers with code
Deep Attention
0 benchmarks · 45 papers with code
Multilingual NLP
0 benchmarks · 42 papers with code
Word Translation
0 benchmarks · 40 papers with code
Morphological Inflection
0 benchmarks · 39 papers with code
text annotation
0 benchmarks · 39 papers with code
Authorship Verification
0 benchmarks · 32 papers with code
Cross-Lingual Word Embeddings
0 benchmarks · 32 papers with code
nlg evaluation
0 benchmarks · 31 papers with code
Hallucination Evaluation
0 benchmarks · 29 papers with code
Text Normalization
0 benchmarks · 27 papers with code
X-ray Classification
0 benchmarks · 27 papers with code
Morphological Tagging
0 benchmarks · 26 papers with code
Comment Generation
0 benchmarks · 23 papers with code
Entity Extraction using GAN
0 benchmarks · 22 papers with code
Review Generation
0 benchmarks · 22 papers with code
Sentence-Pair Classification
0 benchmarks · 22 papers with code
Semantic Composition
0 benchmarks · 21 papers with code
Lexical Simplification
0 benchmarks · 20 papers with code
Punctuation Restoration
0 benchmarks · 19 papers with code
Script Generation
0 benchmarks · 18 papers with code
Text Compression
0 benchmarks · 18 papers with code
Reverse Dictionary
0 benchmarks · 17 papers with code
Diachronic Word Embeddings
0 benchmarks · 14 papers with code
Decipherment
0 benchmarks · 12 papers with code
Pretrained Multilingual Language Models
0 benchmarks · 12 papers with code
Linguistic steganography
0 benchmarks · 11 papers with code
Gender Bias Detection
0 benchmarks · 10 papers with code
Human Agent Collaboration
0 benchmarks · 10 papers with code
Clickbait Detection
0 benchmarks · 9 papers with code
Complex Word Identification
0 benchmarks · 8 papers with code
Sign Language Production
0 benchmarks · 8 papers with code
Text Anonymization
0 benchmarks · 8 papers with code
Toponym Resolution
0 benchmarks · 8 papers with code
Vietnamese Datasets
0 benchmarks · 8 papers with code
Commonsense Causal Reasoning
0 benchmarks · 7 papers with code
Vietnamese Hate Speech Detection
0 benchmarks · 7 papers with code
Abstract Argumentation
0 benchmarks · 6 papers with code
Aggression Identification
0 benchmarks · 6 papers with code
Suggestion mining
0 benchmarks · 6 papers with code
Vietnamese Word Segmentation
0 benchmarks · 6 papers with code
Morphological Disambiguation
0 benchmarks · 5 papers with code
Relevance Detection
0 benchmarks · 5 papers with code
Simultaneous Speech-to-Speech Translation
0 benchmarks · 5 papers with code
Text Attribute Transfer
0 benchmarks · 5 papers with code
AI and Safety
0 benchmarks · 4 papers with code
Music Genre Transfer
0 benchmarks · 4 papers with code
WNLI
0 benchmarks · 4 papers with code
Automated Writing Evaluation
0 benchmarks · 3 papers with code
Cognate Prediction
0 benchmarks · 3 papers with code
Social Media Mental Health Detection
0 benchmarks · 3 papers with code
Text-Variation
0 benchmarks · 3 papers with code
Vietnamese Language Models
0 benchmarks · 3 papers with code
Zero-Shot Machine Translation
0 benchmarks · 3 papers with code
ArabicMMLU
0 benchmarks · 2 papers with code
Author Attribution
0 benchmarks · 2 papers with code
Context Query Reformulation
0 benchmarks · 2 papers with code
Definition Modelling
0 benchmarks · 2 papers with code
Hate Span Identification
0 benchmarks · 2 papers with code
Job Prediction
0 benchmarks · 2 papers with code
Misogynistic Aggression Identification
0 benchmarks · 2 papers with code
News Annotation
0 benchmarks · 2 papers with code
Open Relation Modeling
0 benchmarks · 2 papers with code
Record linking
0 benchmarks · 2 papers with code
Self-Evolving AI
0 benchmarks · 2 papers with code
Syntax Representation
0 benchmarks · 2 papers with code
Text-to-video search
0 benchmarks · 2 papers with code
Turning Point Identification
0 benchmarks · 2 papers with code
Vietnamese Scene Text
0 benchmarks · 2 papers with code
Vietnamese Speech Recognition
0 benchmarks · 2 papers with code
Asynchronous Group Communication
0 benchmarks · 1 paper with code
Coding Problem Tagging
0 benchmarks · 1 paper with code
Collaborative Plan Acquisition
0 benchmarks · 1 paper with code
Cross-lingual Text-to-Image Generation
0 benchmarks · 1 paper with code
Emergent communications on relations
0 benchmarks · 1 paper with code
Emotion Detection and Trigger Summarization
0 benchmarks · 1 paper with code
Extractive Tags Summarization
0 benchmarks · 1 paper with code
Hate Intensity Prediction
0 benchmarks · 1 paper with code
incongruity detection
0 benchmarks · 1 paper with code
Joint Entity and Relation Extraction on Scientific Data
0 benchmarks · 1 paper with code
Joint NER and Classification
0 benchmarks · 1 paper with code
Meme Captioning
0 benchmarks · 1 paper with code
Multi-Dialect Vietnamese
0 benchmarks · 1 paper with code
multi-word expression embedding
0 benchmarks · 1 paper with code
multi-word expression sememe prediction
0 benchmarks · 1 paper with code
Negation and Speculation Scope resolution
0 benchmarks · 1 paper with code
Only Connect Walls Dataset Task 2 (Connections)
0 benchmarks · 1 paper with code
Overlapping Mention Recognition
0 benchmarks · 1 paper with code
Philosophical Reflection
0 benchmarks · 1 paper with code
Phrase Vector Embedding
0 benchmarks · 1 paper with code
Readability optimization
0 benchmarks · 1 paper with code
Reliable Intelligence Identification
0 benchmarks · 1 paper with code
Semi-Supervised Text Regression
0 benchmarks · 1 paper with code
Text Effects Transfer
0 benchmarks · 1 paper with code
Text-to-GQL
0 benchmarks · 1 paper with code
Vietnamese Fact Checking
0 benchmarks · 1 paper with code
Vietnamese Lexical Normalization
0 benchmarks · 1 paper with code
Vietnamese Natural Language Understanding
0 benchmarks · 1 paper with code
Web Page Tagging
0 benchmarks · 1 paper with code
When should a hot water tank be replaced?
0 benchmarks · 1 paper with code
ARQMath2
0 benchmarks · 0 papers with code
Automatic Writing
0 benchmarks · 0 papers with code
Complaint Comment Classification
0 benchmarks · 0 papers with code
Face Selection
0 benchmarks · 0 papers with code
Job classification
0 benchmarks · 0 papers with code
Multi-lingual Text-to-Image Generation
0 benchmarks · 0 papers with code
Multlingual Neural Machine Translation
0 benchmarks · 0 papers with code
Question to Declarative Sentence
0 benchmarks · 0 papers with code
Vietnamese Parsing
0 benchmarks · 0 papers with code
3 tasks in Natural Language Processing are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.
Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.