Browse State-of-the-Art › Natural Language Processing

Natural Language Processing

2,974 benchmarks 943 tasks 3,594 datasets 65,596 papers with code archive 2025-07-28

Syntology code harvested from 16,946 of the papers with code counted above at least one sample ran for 13,242 of them per-sample status is on each paper page

Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 2,998 tables with Natural Language Processing; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).

Parent tasks

190 tasks in Natural Language Processing have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.

Question Answering

142 benchmarks · 4,171 papers with code

Multiple Choice Question Answering (MCQA)

31 benchmarks · 37 papers with code

Zero-Shot Video Question Answer

17 benchmarks · 73 papers with code

Open-Domain Question Answering

15 benchmarks · 238 papers with code

Knowledge Base Question Answering

10 benchmarks · 67 papers with code

Answer Selection

6 benchmarks · 52 papers with code

5 shown of 19 sub-tasks (1 filed under another area). All sub-tasks of Question Answering →

Image Generation

93 benchmarks · 3,102 papers with code

Image-to-Image Translation

38 benchmarks · 550 papers with code

Text-to-Image Generation

17 benchmarks · 546 papers with code

Image Inpainting

12 benchmarks · 331 papers with code

Conditional Image Generation

11 benchmarks · 167 papers with code

Layout-to-Image Generation

11 benchmarks · 24 papers with code

5 shown of 23 sub-tasks (6 filed under another area). All sub-tasks of Image Generation →

Machine Translation

84 benchmarks · 2,444 papers with code

Unsupervised Machine Translation

9 benchmarks · 33 papers with code

Multimodal Machine Translation

3 benchmarks · 38 papers with code

Low-Resource Neural Machine Translation

1 benchmark · 28 papers with code

Transliteration

0 benchmarks · 54 papers with code

Bilingual Lexicon Induction

0 benchmarks · 35 papers with code

5 shown of 10 sub-tasks (1 filed under another area). All sub-tasks of Machine Translation →

80 benchmarks · 974 papers with code

Dynamic Link Prediction

10 benchmarks · 20 papers with code

Inductive Link Prediction

3 benchmarks · 23 papers with code

Link prediction on DH-KGs

1 benchmark · 1 paper with code

Hyperedge Prediction

0 benchmarks · 10 papers with code

Anchor link prediction

0 benchmarks · 1 paper with code

5 shown of 6 sub-tasks (1 filed under another area). All sub-tasks of Link Prediction →

Visual Question Answering (VQA)

76 benchmarks · 1,039 papers with code

Visual Question Answering

29 benchmarks · 1,042 papers with code

Machine Reading Comprehension

4 benchmarks · 206 papers with code

Chart Question Answering

3 benchmarks · 25 papers with code

3D Question Answering (3D-QA)

3 benchmarks · 17 papers with code

Generative Visual Question Answering

1 benchmark · 7 papers with code

5 shown of 9 sub-tasks (2 filed under another area). All sub-tasks of Visual Question Answering (VQA) →

Named Entity Recognition (NER)

76 benchmarks · 955 papers with code

Chinese Named Entity Recognition

7 benchmarks · 40 papers with code

Nested Named Entity Recognition

6 benchmarks · 46 papers with code

Multi-modal Named Entity Recognition

5 benchmarks · 6 papers with code

NER

4 benchmarks · 665 papers with code

Few-shot NER

4 benchmarks · 42 papers with code

5 shown of 13 sub-tasks. All sub-tasks of Named Entity Recognition (NER) →

Text Classification

68 benchmarks · 1,308 papers with code

Document Classification

21 benchmarks · 235 papers with code

Multi-Label Text Classification

20 benchmarks · 78 papers with code

Emotion Classification

9 benchmarks · 117 papers with code

Few-Shot Text Classification

8 benchmarks · 46 papers with code

Topic Models

6 benchmarks · 229 papers with code

5 shown of 16 sub-tasks. All sub-tasks of Text Classification →

Classification

58 benchmarks · 3,778 papers with code

Graph Classification

73 benchmarks · 483 papers with code

Text Classification

68 benchmarks · 1,308 papers with code

Audio Classification

22 benchmarks · 183 papers with code

Medical Image Classification

11 benchmarks · 183 papers with code

Multi-class Classification

5 benchmarks · 289 papers with code

5 shown of 24 sub-tasks (5 filed under another area). All sub-tasks of Classification →

Language Modelling

55 benchmarks · 7,012 papers with code

Long-range modeling

2 benchmarks · 69 papers with code

Cross-Document Language Modeling

2 benchmarks · 1 paper with code

Protein Language Model

1 benchmark · 47 papers with code

XLM-R

0 benchmarks · 99 papers with code

Sentence Pair Modeling

0 benchmarks · 6 papers with code

5 shown of 6 sub-tasks. All sub-tasks of Language Modelling →

Relation Extraction

50 benchmarks · 735 papers with code

Joint Entity and Relation Extraction

16 benchmarks · 56 papers with code

Relation Classification

8 benchmarks · 160 papers with code

Document-level Relation Extraction

4 benchmarks · 66 papers with code

Dialog Relation Extraction

2 benchmarks · 13 papers with code

Relationship Extraction (Distant Supervised)

2 benchmarks · 11 papers with code

5 shown of 15 sub-tasks. All sub-tasks of Relation Extraction →

Sentiment Analysis

42 benchmarks · 1,509 papers with code

Aspect-Based Sentiment Analysis (ABSA)

18 benchmarks · 185 papers with code

Multimodal Sentiment Analysis

5 benchmarks · 93 papers with code

Aspect Sentiment Triplet Extraction

4 benchmarks · 29 papers with code

Aspect Term Extraction and Sentiment Classification

1 benchmark · 8 papers with code

Arabic Sentiment Analysis

1 benchmark · 7 papers with code

5 shown of 12 sub-tasks. All sub-tasks of Sentiment Analysis →

Text Summarization

37 benchmarks · 440 papers with code

Abstractive Text Summarization

15 benchmarks · 362 papers with code

Document Summarization

7 benchmarks · 226 papers with code

Multi-Document Summarization

5 benchmarks · 113 papers with code

Extractive Text Summarization

5 benchmarks · 34 papers with code

Long-Form Narrative Summarization

4 benchmarks · 4 papers with code

5 shown of 12 sub-tasks. All sub-tasks of Text Summarization →

Continual Learning

33 benchmarks · 1,142 papers with code

Class Incremental Learning

6 benchmarks · 296 papers with code

TiROD

1 benchmark · 1 paper with code

Continual Named Entity Recognition

0 benchmarks · 3 papers with code

Continual Panoptic Segmentation

0 benchmarks · 2 papers with code

unsupervised class-incremental learning

0 benchmarks · 2 papers with code

5 shown of 5 sub-tasks (1 filed under another area).

Natural Language Inference

33 benchmarks · 821 papers with code

Cross-Lingual Natural Language Inference

4 benchmarks · 17 papers with code

Visual Entailment

3 benchmarks · 33 papers with code

Answer Generation

2 benchmarks · 111 papers with code

3 shown of 3 sub-tasks.

Image Captioning

33 benchmarks · 774 papers with code

Semi Supervised Learning for Image Captioning

3 benchmarks · 2 papers with code

3D dense captioning

2 benchmarks · 13 papers with code

Hindi Image Captioning

2 benchmarks · 0 papers with code

Relational Captioning

1 benchmark · 2 papers with code

controllable image captioning

0 benchmarks · 8 papers with code

5 shown of 8 sub-tasks (1 filed under another area). All sub-tasks of Image Captioning →

Visual Question Answering

29 benchmarks · 1,042 papers with code

Spatial Reasoning

2 benchmarks · 198 papers with code

Explanatory Visual Question Answering

1 benchmark · 4 papers with code

Object Hallucination

0 benchmarks · 42 papers with code

Vietnamese Visual Question Answering

0 benchmarks · 4 papers with code

MM-Vet v2

0 benchmarks · 1 paper with code

5 shown of 5 sub-tasks (1 filed under another area).

Few-Shot Learning

27 benchmarks · 1,297 papers with code

Few-Shot Semantic Segmentation

13 benchmarks · 102 papers with code

Few-Shot Audio Classification

10 benchmarks · 5 papers with code

Cross-Domain Few-Shot

9 benchmarks · 80 papers with code

Few-Shot Relation Classification

4 benchmarks · 10 papers with code

One-Shot Learning

1 benchmark · 107 papers with code

5 shown of 12 sub-tasks (3 filed under another area). All sub-tasks of Few-Shot Learning →

Code Generation

27 benchmarks · 745 papers with code

Code Documentation Generation

7 benchmarks · 7 papers with code

Code Translation

2 benchmarks · 54 papers with code

Class-level Code Generation

1 benchmark · 4 papers with code

GitHub issue resolution

0 benchmarks · 6 papers with code

Library-Oriented Code Generation

0 benchmarks · 4 papers with code

5 shown of 5 sub-tasks (1 filed under another area).

Data-to-Text Generation

26 benchmarks · 112 papers with code

KG-to-Text Generation

11 benchmarks · 18 papers with code

Unsupervised KG-to-Text Generation

4 benchmarks · 2 papers with code

Visual Storytelling

1 benchmark · 37 papers with code

3 shown of 3 sub-tasks.

Common Sense Reasoning

24 benchmarks · 325 papers with code

Riddle Sense

2 benchmarks · 5 papers with code

Multiview Contextual Commonsense Inference

2 benchmarks · 1 paper with code

Physical Commonsense Reasoning

1 benchmark · 6 papers with code

Discourse Marker Prediction

1 benchmark · 3 papers with code

Empirical Judgments

1 benchmark · 3 papers with code

5 shown of 17 sub-tasks (1 filed under another area). All sub-tasks of Common Sense Reasoning →

Stance Detection

22 benchmarks · 127 papers with code

Stance Detection (US Election 2020 - Biden)

1 benchmark · 1 paper with code

Stance Detection (US Election 2020 - Trump)

1 benchmark · 1 paper with code

Zero-Shot Stance Detection

0 benchmarks · 11 papers with code

Few-Shot Stance Detection

0 benchmarks · 4 papers with code

4 shown of 4 sub-tasks.

Document Classification

21 benchmarks · 235 papers with code

Page Stream Segmentation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Text Generation

20 benchmarks · 2,047 papers with code

Data-to-Text Generation

26 benchmarks · 112 papers with code

Dialogue Generation

13 benchmarks · 265 papers with code

Table-to-Text Generation

8 benchmarks · 43 papers with code

Code Documentation Generation

7 benchmarks · 7 papers with code

Multi-Document Summarization

5 benchmarks · 113 papers with code

5 shown of 25 sub-tasks (1 filed under another area). All sub-tasks of Text Generation →

Semantic Parsing

20 benchmarks · 414 papers with code

Text-To-SQL

10 benchmarks · 209 papers with code

AMR Parsing

8 benchmarks · 51 papers with code

Semantic Dependency Parsing

3 benchmarks · 14 papers with code

DRS Parsing

2 benchmarks · 6 papers with code

UCCA Parsing

2 benchmarks · 4 papers with code

5 shown of 7 sub-tasks. All sub-tasks of Semantic Parsing →

Intent Detection

19 benchmarks · 121 papers with code

Open Intent Detection

17 benchmarks · 6 papers with code

1 shown of 1 sub-task.

Aspect-Based Sentiment Analysis (ABSA)

18 benchmarks · 185 papers with code

Aspect Extraction

6 benchmarks · 34 papers with code

Aspect-Category-Opinion-Sentiment Quadruple Extraction

2 benchmarks · 4 papers with code

Aspect Category Sentiment Analysis

1 benchmark · 12 papers with code

Extract Aspect

1 benchmark · 10 papers with code

Aspect-oriented Opinion Extraction

1 benchmark · 4 papers with code

5 shown of 8 sub-tasks. All sub-tasks of Aspect-Based Sentiment Analysis (ABSA) →

Text-to-Image Generation

17 benchmarks · 546 papers with code

Conditional Text-to-Image Synthesis

3 benchmarks · 8 papers with code

Text-based Image Editing

1 benchmark · 29 papers with code

text-guided-image-editing

0 benchmarks · 34 papers with code

Concept Alignment

0 benchmarks · 19 papers with code

Zero-Shot Text-to-Image Generation

0 benchmarks · 11 papers with code

5 shown of 7 sub-tasks. All sub-tasks of Text-to-Image Generation →

Video Generation

16 benchmarks · 609 papers with code

Unconditional Video Generation

1 benchmark · 10 papers with code

Image to Video Generation

0 benchmarks · 38 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Prompt Engineering

16 benchmarks · 454 papers with code

Visual Prompting

0 benchmarks · 60 papers with code

1 shown of 1 sub-task.

Dependency Parsing

16 benchmarks · 338 papers with code

Dependency Grammar Induction

2 benchmarks · 3 papers with code

Unsupervised Dependency Parsing

1 benchmark · 5 papers with code

Cross-lingual zero-shot dependency parsing

1 benchmark · 3 papers with code

Transition-Based Dependency Parsing

0 benchmarks · 12 papers with code

Prepositional Phrase Attachment

0 benchmarks · 5 papers with code

5 shown of 5 sub-tasks.

Coreference Resolution

16 benchmarks · 288 papers with code

coreference-resolution

0 benchmarks · 215 papers with code

Cross Document Coreference Resolution

0 benchmarks · 10 papers with code

2 shown of 2 sub-tasks.

Word Sense Disambiguation

16 benchmarks · 150 papers with code

Word Sense Induction

1 benchmark · 19 papers with code

1 shown of 1 sub-task.

Abstractive Text Summarization

15 benchmarks · 362 papers with code

Timeline Summarization

1 benchmark · 8 papers with code

Multimodal Abstractive Text Summarization

1 benchmark · 2 papers with code

Reader-Aware Summarization

1 benchmark · 0 papers with code

3 shown of 3 sub-tasks.

Part-Of-Speech Tagging

15 benchmarks · 228 papers with code

Unsupervised Part-Of-Speech Tagging

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Hate Speech Detection

15 benchmarks · 203 papers with code

Hope Speech Detection

2 benchmarks · 9 papers with code

Hate Speech Normalization

0 benchmarks · 3 papers with code

Hate Speech Detection CrisisHateMM Benchmark

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Slot Filling

14 benchmarks · 140 papers with code

Zero-shot Slot Filling

3 benchmarks · 10 papers with code

Extracting COVID-19 Events from Twitter

1 benchmark · 2 papers with code

2 shown of 2 sub-tasks.

Semantic Textual Similarity

13 benchmarks · 693 papers with code

Paraphrase Identification

11 benchmarks · 76 papers with code

Cross-Lingual Semantic Textual Similarity

0 benchmarks · 3 papers with code

2 shown of 2 sub-tasks.

Dialogue Generation

13 benchmarks · 265 papers with code

Multi-modal Dialogue Generation

1 benchmark · 3 papers with code

1 shown of 1 sub-task.

Cross-Modal Retrieval

13 benchmarks · 244 papers with code

Cross-modal retrieval with noisy correspondence

3 benchmarks · 14 papers with code

Image-text matching

1 benchmark · 102 papers with code

Zero-shot Composed Person Retrieval

1 benchmark · 1 paper with code

multilingual cross-modal retrieval

0 benchmarks · 2 papers with code

Cross-Modal Retrieval on RSITMD

0 benchmarks · 0 papers with code

5 shown of 5 sub-tasks.

Grammatical Error Correction

13 benchmarks · 142 papers with code

Grammatical Error Detection

4 benchmarks · 19 papers with code

1 shown of 1 sub-task.

Open Information Extraction

13 benchmarks · 67 papers with code

Event Extraction

9 benchmarks · 149 papers with code

1 shown of 1 sub-task.

Retrieval

11 benchmarks · 5,274 papers with code

Text Retrieval

16 benchmarks · 335 papers with code

Table Retrieval

1 benchmark · 15 papers with code

Deep Hashing

0 benchmarks · 56 papers with code

3 shown of 3 sub-tasks.

Information Retrieval

11 benchmarks · 1,188 papers with code

Passage Retrieval

5 benchmarks · 133 papers with code

Scientific Results Extraction

2 benchmarks · 5 papers with code

Zero Shot on BEIR (Inference Free Model)

1 benchmark · 4 papers with code

TAR

0 benchmarks · 35 papers with code

Cross-Lingual Information Retrieval

0 benchmarks · 14 papers with code

5 shown of 6 sub-tasks. All sub-tasks of Information Retrieval →

Question Generation

11 benchmarks · 264 papers with code

Poll Generation

1 benchmark · 3 papers with code

1 shown of 1 sub-task.

Speech-to-Text Translation

11 benchmarks · 64 papers with code

Simultaneous Speech-to-Text Translation

0 benchmarks · 3 papers with code

1 shown of 1 sub-task.

Entity Resolution

11 benchmarks · 55 papers with code

Blocking

5 benchmarks · 129 papers with code

1 shown of 1 sub-task.

2D Semantic Segmentation

11 benchmarks · 49 papers with code

Image Segmentation

13 benchmarks · 2,073 papers with code

Human Part Segmentation

6 benchmarks · 15 papers with code

Reflection Removal

5 benchmarks · 38 papers with code

Continual Semantic Segmentation

3 benchmarks · 17 papers with code

Text Style Transfer

2 benchmarks · 93 papers with code

5 shown of 17 sub-tasks (5 filed under another area). All sub-tasks of 2D Semantic Segmentation →

Mathematical Reasoning

10 benchmarks · 395 papers with code

Math Word Problem Solving

13 benchmarks · 80 papers with code

Formal Logic

1 benchmark · 20 papers with code

Abstract Algebra

1 benchmark · 6 papers with code

Mathematical Induction

1 benchmark · 4 papers with code

High School Mathematics

1 benchmark · 1 paper with code

5 shown of 7 sub-tasks (2 filed under another area). All sub-tasks of Mathematical Reasoning →

Text-To-SQL

10 benchmarks · 209 papers with code

MMSQL performance

1 benchmark · 1 paper with code

1 shown of 1 sub-task.

Entity Alignment

10 benchmarks · 117 papers with code

Multi-modal Entity Alignment

8 benchmarks · 14 papers with code

1 shown of 1 sub-task (1 filed under another area).

Cross-Lingual Document Classification

10 benchmarks · 12 papers with code

News Classification

4 benchmarks · 33 papers with code

1 shown of 1 sub-task.

Emotion Recognition

9 benchmarks · 614 papers with code

Speech Emotion Recognition

16 benchmarks · 139 papers with code

Emotion Recognition in Conversation

16 benchmarks · 83 papers with code

Multimodal Emotion Recognition

7 benchmarks · 80 papers with code

Emotion Recognition in Context

4 benchmarks · 5 papers with code

EEG Emotion Recognition

3 benchmarks · 14 papers with code

5 shown of 12 sub-tasks (2 filed under another area). All sub-tasks of Emotion Recognition →

Event Extraction

9 benchmarks · 149 papers with code

NER

4 benchmarks · 665 papers with code

Event Causality Identification

0 benchmarks · 13 papers with code

Zero-shot Event Extraction

0 benchmarks · 4 papers with code

3 shown of 3 sub-tasks.

Source Code Summarization

9 benchmarks · 39 papers with code

Method name prediction

1 benchmark · 15 papers with code

1 shown of 1 sub-task.

Relation Classification

8 benchmarks · 160 papers with code

Few-Shot Relation Classification

4 benchmarks · 10 papers with code

Implicit Discourse Relation Classification

0 benchmarks · 7 papers with code

Cause-Effect Relation Classification

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Entity Typing

8 benchmarks · 92 papers with code

Entity Typing on DH-KGs

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Few-Shot Text Classification

8 benchmarks · 46 papers with code

Zero-Shot Out-of-Domain Detection

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Table-to-Text Generation

8 benchmarks · 43 papers with code

KB-to-Language Generation

1 benchmark · 3 papers with code

1 shown of 1 sub-task.

Knowledge Distillation

7 benchmarks · 1,740 papers with code

Data-free Knowledge Distillation

2 benchmarks · 37 papers with code

Self-Knowledge Distillation

0 benchmarks · 36 papers with code

2 shown of 2 sub-tasks.

Reading Comprehension

7 benchmarks · 634 papers with code

Machine Reading Comprehension

4 benchmarks · 206 papers with code

Intent Recognition

1 benchmark · 22 papers with code

Implicit Relations

1 benchmark · 20 papers with code

Question Selection

1 benchmark · 18 papers with code

LAMBADA

1 benchmark · 15 papers with code

5 shown of 20 sub-tasks (1 filed under another area). All sub-tasks of Reading Comprehension →

Semantic Similarity

7 benchmarks · 534 papers with code

Semantic Shift Detection

0 benchmarks · 5 papers with code

Similarity Explanation

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Knowledge Graph Completion

7 benchmarks · 251 papers with code

Inductive knowledge graph completion

3 benchmarks · 14 papers with code

Triple Classification

1 benchmark · 23 papers with code

Large Language Model

0 benchmarks · 2,250 papers with code

Inductive Relation Prediction

0 benchmarks · 11 papers with code

4 shown of 4 sub-tasks (1 filed under another area).

Document Summarization

7 benchmarks · 226 papers with code

Email Thread Summarization

2 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Semantic Role Labeling

7 benchmarks · 138 papers with code

Predicate Detection

3 benchmarks · 9 papers with code

Semantic Role Labeling (predicted predicates)

2 benchmarks · 4 papers with code

Textual Analogy Parsing

0 benchmarks · 2 papers with code

3 shown of 3 sub-tasks.

Binary text classification

7 benchmarks · 12 papers with code

Detection of potentially void clauses

1 benchmark · 1 paper with code

1 shown of 1 sub-task.

Natural Language Understanding

6 benchmarks · 809 papers with code

Vietnamese Social Media Text Processing

0 benchmarks · 9 papers with code

Emotional Dialogue Acts

0 benchmarks · 2 papers with code

2 shown of 2 sub-tasks.

Optical Character Recognition (OCR)

6 benchmarks · 462 papers with code

Handwritten Text Recognition

13 benchmarks · 56 papers with code

Handwriting Recognition

3 benchmarks · 53 papers with code

Handwritten Digit Recognition

2 benchmarks · 29 papers with code

Active Learning

1 benchmark · 913 papers with code

Irregular Text Recognition

0 benchmarks · 5 papers with code

5 shown of 10 sub-tasks (2 filed under another area). All sub-tasks of Optical Character Recognition (OCR) →

Topic Models

6 benchmarks · 229 papers with code

Topic coverage

3 benchmarks · 6 papers with code

Dynamic Topic Modeling

0 benchmarks · 3 papers with code

2 shown of 2 sub-tasks.

Intrusion Detection

6 benchmarks · 151 papers with code

Network Intrusion Detection

6 benchmarks · 67 papers with code

1 shown of 1 sub-task.

Language Identification

6 benchmarks · 143 papers with code

Native Language Identification

1 benchmark · 5 papers with code

Dialect Identification

0 benchmarks · 33 papers with code

2 shown of 2 sub-tasks.

Sentence Classification

6 benchmarks · 115 papers with code

Unfairness Detection

0 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Text-To-Speech Synthesis

6 benchmarks · 104 papers with code

Prosody Prediction

1 benchmark · 4 papers with code

Zero-Shot Multi-Speaker TTS

0 benchmarks · 3 papers with code

2 shown of 2 sub-tasks.

Text-to-Video Generation

6 benchmarks · 97 papers with code

Text-to-Video Editing

0 benchmarks · 6 papers with code

Subject-driven Video Generation

0 benchmarks · 2 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Key Information Extraction

6 benchmarks · 40 papers with code

Key-value Pair Extraction

2 benchmarks · 9 papers with code

1 shown of 1 sub-task.

Aspect Extraction

6 benchmarks · 34 papers with code

Hidden Aspect Detection

0 benchmarks · 2 papers with code

Latent Aspect Detection

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Representation Learning

5 benchmarks · 4,662 papers with code

Disentanglement

3 benchmarks · 744 papers with code

Graph Representation Learning

1 benchmark · 479 papers with code

Feature Upsampling

1 benchmark · 22 papers with code

Sentence Embeddings

0 benchmarks · 255 papers with code

Network Embedding

0 benchmarks · 165 papers with code

5 shown of 14 sub-tasks (2 filed under another area). All sub-tasks of Representation Learning →

Deep Clustering

5 benchmarks · 135 papers with code

Trajectory Clustering

0 benchmarks · 6 papers with code

Deep Nonparametric Clustering

0 benchmarks · 1 paper with code

NONPARAMETRIC DEEP CLUSTERING

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Story Generation

5 benchmarks · 110 papers with code

Visual Storytelling

1 benchmark · 37 papers with code

1 shown of 1 sub-task.

Bias Detection

5 benchmarks · 80 papers with code

Selection bias

0 benchmarks · 143 papers with code

1 shown of 1 sub-task.

Phrase Grounding

5 benchmarks · 51 papers with code

Grounded Open Vocabulary Acquisition

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Extractive Text Summarization

5 benchmarks · 34 papers with code

Reader-Aware Summarization

1 benchmark · 0 papers with code

1 shown of 1 sub-task.

Binary Classification

4 benchmarks · 710 papers with code

Cancer-no cancer per breast classification

2 benchmarks · 2 papers with code

Cancer-no cancer per image classification

1 benchmark · 2 papers with code

Stable MCI vs Progressive MCI

1 benchmark · 1 paper with code

LLM-generated Text Detection

0 benchmarks · 11 papers with code

5 shown of 6 sub-tasks. All sub-tasks of Binary Classification →

Task-Oriented Dialogue Systems

4 benchmarks · 131 papers with code

SSTOD

2 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Constituency Parsing

4 benchmarks · 81 papers with code

Constituency Grammar Induction

4 benchmarks · 20 papers with code

1 shown of 1 sub-task.

Document-level Relation Extraction

4 benchmarks · 66 papers with code

Document-level RE with incomplete labeling

2 benchmarks · 3 papers with code

1 shown of 1 sub-task.

Discourse Parsing

4 benchmarks · 57 papers with code

End-to-End RST Parsing

1 benchmark · 3 papers with code

Discourse Segmentation

0 benchmarks · 14 papers with code

Connective Detection

0 benchmarks · 2 papers with code

3 shown of 3 sub-tasks.

Document Text Classification

4 benchmarks · 5 papers with code

Learning with noisy labels

20 benchmarks · 143 papers with code

Multi-Label Classification Of Biomedical Texts

2 benchmarks · 4 papers with code

Political Salient Issue Orientation Detection

1 benchmark · 3 papers with code

3 shown of 3 sub-tasks (1 filed under another area).

Data Augmentation

3 benchmarks · 3,225 papers with code

Image Augmentation

1 benchmark · 127 papers with code

Text Augmentation

0 benchmarks · 41 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Paraphrase Generation

3 benchmarks · 74 papers with code

Multilingual Paraphrase Generation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

3D Action Recognition

3 benchmarks · 38 papers with code

Skeleton Based Action Recognition

34 benchmarks · 219 papers with code

Image Manipulation Detection

21 benchmarks · 38 papers with code

Zero Shot Skeletal Action Recognition

3 benchmarks · 7 papers with code

Generalized Zero Shot skeletal action recognition

3 benchmarks · 3 papers with code

Model Editing

0 benchmarks · 107 papers with code

5 shown of 6 sub-tasks (2 filed under another area). All sub-tasks of 3D Action Recognition →

Multimodal Machine Translation

3 benchmarks · 38 papers with code

Multimodal Lexical Translation

4 benchmarks · 2 papers with code

Face to Face Translation

0 benchmarks · 4 papers with code

2 shown of 2 sub-tasks.

Text Clustering

3 benchmarks · 38 papers with code

Short Text Clustering

8 benchmarks · 18 papers with code

Open Intent Discovery

6 benchmarks · 5 papers with code

Hierarchical Text Clustering

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Meme Classification

3 benchmarks · 30 papers with code

Hateful Meme Classification

4 benchmarks · 9 papers with code

1 shown of 1 sub-task.

Age And Gender Classification

3 benchmarks · 16 papers with code

Author Profiling

0 benchmarks · 14 papers with code

1 shown of 1 sub-task.

Inductive knowledge graph completion

3 benchmarks · 14 papers with code

Large Language Model

0 benchmarks · 2,250 papers with code

1 shown of 1 sub-task.

Biomedical Information Retrieval

3 benchmarks · 7 papers with code

SpO2 estimation

3 benchmarks · 3 papers with code

PICO

1 benchmark · 26 papers with code

2 shown of 2 sub-tasks.

Multimodal Text and Image Classification

3 benchmarks · 5 papers with code

image-sentence alignment

12 benchmarks · 2 papers with code

Open-World Social Event Classification

0 benchmarks · 0 papers with code

2 shown of 2 sub-tasks.

Text Style Transfer

2 benchmarks · 93 papers with code

Formality Style Transfer

1 benchmark · 10 papers with code

Semi-Supervised Formality Style Transfer

0 benchmarks · 1 paper with code

Word Attribute Transfer

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Document Ranking

2 benchmarks · 65 papers with code

Session Search

0 benchmarks · 2 papers with code

1 shown of 1 sub-task.

Term Extraction

2 benchmarks · 44 papers with code

Nested Term Extraction

3 benchmarks · 2 papers with code

Nested Term Recognition

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Data-free Knowledge Distillation

2 benchmarks · 37 papers with code

Benchmarking

2 benchmarks · 2,658 papers with code

1 shown of 1 sub-task (1 filed under another area).

Weakly Supervised Classification

2 benchmarks · 25 papers with code

Weakly Supervised Data Denoising

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Temporal Information Extraction

2 benchmarks · 20 papers with code

Temporal Tagging

8 benchmarks · 4 papers with code

1 shown of 1 sub-task.

Hope Speech Detection

2 benchmarks · 9 papers with code

Hope Speech Detection for English

1 benchmark · 1 paper with code

Hope Speech Detection for Malayalam

1 benchmark · 1 paper with code

Hope Speech Detection for Tamil

1 benchmark · 1 paper with code

3 shown of 3 sub-tasks.

Handwriting Verification

2 benchmarks · 6 papers with code

Bangla Spelling Error Correction

1 benchmark · 3 papers with code

1 shown of 1 sub-task.

Aspect-Category-Opinion-Sentiment Quadruple Extraction

2 benchmarks · 4 papers with code

Conversational Sentiment Quadruple Extraction

2 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Recognizing Emotion Cause in Conversations

2 benchmarks · 2 papers with code

Causal Emotion Entailment

1 benchmark · 10 papers with code

1 shown of 1 sub-task.

Reinforcement Learning

1 benchmark · 4,183 papers with code

Deep Reinforcement Learning

0 benchmarks · 1,739 papers with code

1 shown of 1 sub-task.

Active Learning

1 benchmark · 913 papers with code

Active Object Detection

2 benchmarks · 6 papers with code

1 shown of 1 sub-task.

Instruction Following

1 benchmark · 609 papers with code

visual instruction following

1 benchmark · 14 papers with code

1 shown of 1 sub-task.

Cross-Lingual Transfer

1 benchmark · 333 papers with code

Cross-Lingual NER

28 benchmarks · 27 papers with code

Zero-Shot Cross-Lingual Transfer

2 benchmarks · 80 papers with code

2 shown of 2 sub-tasks.

Chatbot

1 benchmark · 269 papers with code

Dialogue Generation

13 benchmarks · 265 papers with code

1 shown of 1 sub-task.

Natural Language Queries

1 benchmark · 135 papers with code

Text-to-CQL

0 benchmarks · 0 papers with code

1 shown of 1 sub-task.

Argument Mining

1 benchmark · 98 papers with code

ValNov

2 benchmarks · 1 paper with code

Component Classification

1 benchmark · 8 papers with code

Argument Pair Extraction (APE)

1 benchmark · 4 papers with code

Claim Extraction with Stance Classification (CESC)

1 benchmark · 1 paper with code

Claim-Evidence Pair Extraction (CEPE)

1 benchmark · 1 paper with code

5 shown of 6 sub-tasks. All sub-tasks of Argument Mining →

Multimodal Deep Learning

1 benchmark · 97 papers with code

Multimodal Text and Image Classification

3 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Language Acquisition

1 benchmark · 83 papers with code

Grounded language learning

0 benchmarks · 23 papers with code

1 shown of 1 sub-task.

Conversational Question Answering

1 benchmark · 63 papers with code

Question Rewriting

0 benchmarks · 15 papers with code

1 shown of 1 sub-task.

Sentence Completion

1 benchmark · 49 papers with code

Hurtful Sentence Completion

1 benchmark · 1 paper with code

1 shown of 1 sub-task.

Spam detection

1 benchmark · 38 papers with code

Traditional Spam Detection

1 benchmark · 2 papers with code

Context-specific Spam Detection

1 benchmark · 1 paper with code

2 shown of 2 sub-tasks.

Visual Storytelling

1 benchmark · 37 papers with code

Image-guided Story Ending Generation

2 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Conditional Text Generation

1 benchmark · 32 papers with code

Contextualized Literature-based Discovery

0 benchmarks · 1 paper with code

Multimedia Generative Script Learning

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Open-Domain Dialog

1 benchmark · 32 papers with code

Dialogue Evaluation

2 benchmarks · 61 papers with code

1 shown of 1 sub-task.

Temporal Relation Extraction

1 benchmark · 32 papers with code

Temporal Relation Classification

4 benchmarks · 7 papers with code

1 shown of 1 sub-task.

Ad-Hoc Information Retrieval

1 benchmark · 27 papers with code

Document Ranking

2 benchmarks · 65 papers with code

1 shown of 1 sub-task.

Document AI

1 benchmark · 24 papers with code

document understanding

0 benchmarks · 140 papers with code

1 shown of 1 sub-task.

Goal-Oriented Dialog

1 benchmark · 24 papers with code

User Simulation

0 benchmarks · 25 papers with code

1 shown of 1 sub-task.

Sentence Compression

1 benchmark · 22 papers with code

Unsupervised Sentence Compression

0 benchmarks · 4 papers with code

1 shown of 1 sub-task.

Emotional Intelligence

1 benchmark · 19 papers with code

Dark Humor Detection

1 benchmark · 2 papers with code

SNARKS

0 benchmarks · 6 papers with code

Ruin Names

0 benchmarks · 3 papers with code

3 shown of 3 sub-tasks.

Question Similarity

1 benchmark · 17 papers with code

Medical question pair similarity computation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Lexical Normalization

1 benchmark · 16 papers with code

Pronunciation Dictionary Creation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Attribute Extraction

1 benchmark · 11 papers with code

legal outcome extraction

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Clinical Concept Extraction

1 benchmark · 9 papers with code

Clinical Information Retreival

0 benchmarks · 2 papers with code

1 shown of 1 sub-task.

Scientific Document Summarization

1 benchmark · 9 papers with code

Lay Summarization

2 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Summarization

1 benchmark · 4 papers with code

Unsupervised Extractive Summarization

3 benchmarks · 18 papers with code

Query-focused Summarization

0 benchmarks · 18 papers with code

2 shown of 2 sub-tasks.

1 Image, 2*2 Stitchi

1 benchmark · 3 papers with code

Pose Estimation

31 benchmarks · 1,679 papers with code

Text-to-Image Generation

17 benchmarks · 546 papers with code

Image Deblurring

9 benchmarks · 167 papers with code

Virtual Try-on

9 benchmarks · 114 papers with code

Style Transfer

3 benchmarks · 759 papers with code

5 shown of 12 sub-tasks (9 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →

Information Extraction

1 benchmark · 3 papers with code

Joint Entity and Relation Extraction

16 benchmarks · 56 papers with code

Event Extraction

9 benchmarks · 149 papers with code

Attribute Value Extraction

4 benchmarks · 16 papers with code

Low Resource Named Entity Recognition

3 benchmarks · 15 papers with code

Drug–drug Interaction Extraction

3 benchmarks · 13 papers with code

5 shown of 16 sub-tasks (1 filed under another area). All sub-tasks of Information Extraction →

Text2text Generation

1 benchmark · 2 papers with code

Keyphrase Generation

0 benchmarks · 43 papers with code

Figurative Language Visualization

0 benchmarks · 1 paper with code

Sketch-to-text Generation

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Dialogue

1 benchmark · 1 paper with code

Dialogue Generation

13 benchmarks · 265 papers with code

Visual Dialog

8 benchmarks · 56 papers with code

Dialogue State Tracking

7 benchmarks · 138 papers with code

Dialogue Act Classification

5 benchmarks · 23 papers with code

Task-Oriented Dialogue Systems

4 benchmarks · 131 papers with code

5 shown of 21 sub-tasks (1 filed under another area). All sub-tasks of Dialogue →

Language Modeling

0 benchmarks · 5,620 papers with code

Dream Generation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Deep Learning

0 benchmarks · 2,693 papers with code

Polynomial Neural Networks

0 benchmarks · 10 papers with code

1 shown of 1 sub-task.

Large Language Model

0 benchmarks · 2,250 papers with code

Knowledge Graphs

4 benchmarks · 1,273 papers with code

RAG

1 benchmark · 758 papers with code

AI Agent

0 benchmarks · 111 papers with code

3 shown of 3 sub-tasks (2 filed under another area).

Word Embeddings

0 benchmarks · 1,177 papers with code

Learning Word Embeddings

0 benchmarks · 24 papers with code

Multilingual Word Embeddings

0 benchmarks · 19 papers with code

Embeddings Evaluation

0 benchmarks · 10 papers with code

Contextualised Word Representations

0 benchmarks · 4 papers with code

4 shown of 4 sub-tasks.

NMT

0 benchmarks · 523 papers with code

Direct NMT

0 benchmarks · 0 papers with code

1 shown of 1 sub-task.

Text to Image Generation

0 benchmarks · 461 papers with code

Text to 3D

1 benchmark · 102 papers with code

1 shown of 1 sub-task (1 filed under another area).

Sentence Embeddings

0 benchmarks · 255 papers with code

Sentence Embeddings For Biomedical Texts

2 benchmarks · 2 papers with code

Sentence Compression

1 benchmark · 22 papers with code

Sentence Embedding

0 benchmarks · 162 papers with code

Joint Multilingual Sentence Representations

0 benchmarks · 2 papers with code

4 shown of 4 sub-tasks.

Symbolic Regression

0 benchmarks · 155 papers with code

Equation Discovery

0 benchmarks · 34 papers with code

1 shown of 1 sub-task.

document understanding

0 benchmarks · 140 papers with code

Line Items Extraction

0 benchmarks · 0 papers with code

1 shown of 1 sub-task.

Data Integration

0 benchmarks · 130 papers with code

Entity Resolution

11 benchmarks · 55 papers with code

Entity Alignment

10 benchmarks · 117 papers with code

Table annotation

0 benchmarks · 23 papers with code

3 shown of 3 sub-tasks.

Model Editing

0 benchmarks · 107 papers with code

knowledge editing

1 benchmark · 88 papers with code

1 shown of 1 sub-task.

Authorship Attribution

0 benchmarks · 62 papers with code

Source Code Authorship Attribution

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

De-identification

0 benchmarks · 54 papers with code

Privacy Preserving Deep Learning

0 benchmarks · 30 papers with code

Full-body anonymization

0 benchmarks · 4 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Spelling Correction

0 benchmarks · 52 papers with code

Bangla Spelling Error Correction

1 benchmark · 3 papers with code

1 shown of 1 sub-task.

Token Classification

0 benchmarks · 47 papers with code

Blackout Poetry Generation

1 benchmark · 1 paper with code

Toxic Spans Detection

0 benchmarks · 18 papers with code

2 shown of 2 sub-tasks.

Dialogue Understanding

0 benchmarks · 35 papers with code

Spoken Language Understanding

5 benchmarks · 135 papers with code

Dialogue Safety Prediction

2 benchmarks · 3 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Abuse Detection

0 benchmarks · 32 papers with code

Hate Speech Detection

15 benchmarks · 203 papers with code

1 shown of 1 sub-task.

Constrained Clustering

0 benchmarks · 29 papers with code

Incremental Constrained Clustering

2 benchmarks · 1 paper with code

Only Connect Walls Dataset Task 1 (Grouping)

1 benchmark · 10 papers with code

2 shown of 2 sub-tasks.

Stock Prediction

0 benchmarks · 29 papers with code

PAIR TRADING

2 benchmarks · 2 papers with code

Text-Based Stock Prediction

0 benchmarks · 4 papers with code

Event-Driven Trading

0 benchmarks · 1 paper with code

3 shown of 3 sub-tasks.

Table annotation

0 benchmarks · 23 papers with code

Column Type Annotation

12 benchmarks · 19 papers with code

Cell Entity Annotation

5 benchmarks · 6 papers with code

Columns Property Annotation

4 benchmarks · 5 papers with code

Row Annotation

1 benchmark · 1 paper with code

Table Type Detection

1 benchmark · 1 paper with code

5 shown of 6 sub-tasks. All sub-tasks of Table annotation →

Sentence Summarization

0 benchmarks · 19 papers with code

Unsupervised Sentence Summarization

0 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Conversational Response Generation

0 benchmarks · 17 papers with code

Personalized and Emotional Conversation

1 benchmark · 1 paper with code

1 shown of 1 sub-task.

Propaganda detection

0 benchmarks · 15 papers with code

Propaganda span identification

0 benchmarks · 1 paper with code

Propaganda technique identification

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Twitter Sentiment Analysis

0 benchmarks · 14 papers with code

Tweet-Reply Sentiment Analysis

1 benchmark · 0 papers with code

1 shown of 1 sub-task.

News Generation

0 benchmarks · 12 papers with code

Headline Generation

1 benchmark · 36 papers with code

1 shown of 1 sub-task.

Negation Detection

0 benchmarks · 11 papers with code

Negation Scope Resolution

4 benchmarks · 4 papers with code

1 shown of 1 sub-task.

Cross-Lingual Entity Linking

0 benchmarks · 9 papers with code

Variable Disambiguation

1 benchmark · 1 paper with code

1 shown of 1 sub-task.

Lexical Analysis

0 benchmarks · 8 papers with code

Lexical Complexity Prediction

0 benchmarks · 11 papers with code

1 shown of 1 sub-task.

Literature Mining

0 benchmarks · 7 papers with code

Systematic Literature Review

0 benchmarks · 49 papers with code

1 shown of 1 sub-task.

Persian Sentiment Analysis

0 benchmarks · 7 papers with code

Transition-Based Dependency Parsing

0 benchmarks · 12 papers with code

1 shown of 1 sub-task.

Sentence Pair Modeling

0 benchmarks · 6 papers with code

Semantic Similarity

7 benchmarks · 534 papers with code

1 shown of 1 sub-task.

Data Mining

0 benchmarks · 4 papers with code

Argument Mining

1 benchmark · 98 papers with code

Opinion Mining

1 benchmark · 64 papers with code

Sequential Pattern Mining

1 benchmark · 9 papers with code

CSV dialect detection

1 benchmark · 1 paper with code

cognitive diagnosis

0 benchmarks · 22 papers with code

5 shown of 7 sub-tasks. All sub-tasks of Data Mining →

Continual Named Entity Recognition

0 benchmarks · 3 papers with code

FG-1-PG-1

3 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Multimodal Association

0 benchmarks · 3 papers with code

multimodal generation

1 benchmark · 53 papers with code

1 shown of 1 sub-task.

Speculation Detection

0 benchmarks · 3 papers with code

Speculation Scope Resolution

3 benchmarks · 2 papers with code

1 shown of 1 sub-task.

Joint Multilingual Sentence Representations

0 benchmarks · 2 papers with code

Abstract Meaning Representation

0 benchmarks · 92 papers with code

1 shown of 1 sub-task.

Natural Language Transduction

0 benchmarks · 2 papers with code

Lipreading

8 benchmarks · 36 papers with code

1 shown of 1 sub-task (1 filed under another area).

Personality Generation

0 benchmarks · 2 papers with code

Personality Alignment

0 benchmarks · 3 papers with code

1 shown of 1 sub-task.

Political evalutation

0 benchmarks · 2 papers with code

Alignement visualisation

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

trustable and focussed LLM generated content

0 benchmarks · 2 papers with code

Game Design

0 benchmarks · 14 papers with code

1 shown of 1 sub-task.

Vietnamese Aspect-Based Sentiment Analysis

0 benchmarks · 2 papers with code

Sentiment Dependency Learning

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Vietnamese Sentiment Analysis

0 benchmarks · 2 papers with code

Vietnamese Multimodal Sentiment Analysis

0 benchmarks · 1 paper with code

1 shown of 1 sub-task.

Nested Term Recognition

0 benchmarks · 1 paper with code

Nested Term Recognition from Flat Supervision

3 benchmarks · 1 paper with code

1 shown of 1 sub-task.

NLP based Person Retrival

0 benchmarks · 1 paper with code

Decoder

1 benchmark · 4,358 papers with code

1 shown of 1 sub-task.

Anaphora Resolution

0 benchmarks · 0 papers with code

Abstract Anaphora Resolution

1 benchmark · 1 paper with code

Bridging Anaphora Resolution

0 benchmarks · 2 papers with code

2 shown of 2 sub-tasks.

Chinese

0 benchmarks · 0 papers with code

Chinese Word Segmentation

6 benchmarks · 50 papers with code

Handwritten Chinese Text Recognition

0 benchmarks · 3 papers with code

Chinese Spelling Error Correction

0 benchmarks · 2 papers with code

Chinese Zero Pronoun Resolution

0 benchmarks · 2 papers with code

Offline Handwritten Chinese Character Recognition

0 benchmarks · 2 papers with code

5 shown of 5 sub-tasks.

Cross-Lingual

0 benchmarks · 0 papers with code

Cross-Lingual Document Classification

10 benchmarks · 12 papers with code

Cross-Lingual Transfer

1 benchmark · 333 papers with code

Cross-Lingual Entity Linking

0 benchmarks · 9 papers with code

Cross-Language Text Summarization

0 benchmarks · 0 papers with code

4 shown of 4 sub-tasks.

Optical Charater Recogntion

0 benchmarks · 0 papers with code

Bangla Text Detection

1 benchmark · 0 papers with code

1 shown of 1 sub-task.

Pcl Detection

0 benchmarks · 0 papers with code

SemEval-2022 Task 4-1 (Binary PCL Detection)

1 benchmark · 0 papers with code

SemEval-2022 Task 4-2 (Multi-label PCL Detection)

1 benchmark · 0 papers with code

2 shown of 2 sub-tasks.

Shallow Syntax

0 benchmarks · 0 papers with code

Chunking

5 benchmarks · 120 papers with code

1 shown of 1 sub-task.

Taxonomy Learning

0 benchmarks · 0 papers with code

Hypernym Discovery

3 benchmarks · 8 papers with code

Taxonomy Expansion

0 benchmarks · 13 papers with code

2 shown of 2 sub-tasks.

Temporal Processing

0 benchmarks · 0 papers with code

Temporal Information Extraction

2 benchmarks · 20 papers with code

Timex normalization

2 benchmarks · 5 papers with code

Document Dating

2 benchmarks · 2 papers with code

3 shown of 3 sub-tasks.

Tasks with no parent task

224 tasks in Natural Language Processing sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.

Entity Linking

27 benchmarks · 277 papers with code

Conversational Response Selection

14 benchmarks · 36 papers with code

Text Simplification

11 benchmarks · 129 papers with code

Entity Disambiguation

11 benchmarks · 63 papers with code

Fake News Detection

10 benchmarks · 203 papers with code

Sarcasm Detection

9 benchmarks · 76 papers with code

Image-to-Text Retrieval

8 benchmarks · 37 papers with code

Word Alignment

7 benchmarks · 92 papers with code

Explanation Generation

5 benchmarks · 92 papers with code

Long-Context Understanding

5 benchmarks · 54 papers with code

Keyphrase Extraction

5 benchmarks · 52 papers with code

Linguistic Acceptability

5 benchmarks · 49 papers with code

Recipe Generation

5 benchmarks · 15 papers with code

Intent Classification

4 benchmarks · 113 papers with code

Aspect Category Detection

4 benchmarks · 13 papers with code

Extreme Summarization

4 benchmarks · 13 papers with code

Cross-Lingual Bitext Mining

4 benchmarks · 5 papers with code

Translation

3 benchmarks · 3,574 papers with code

Response Generation

3 benchmarks · 370 papers with code

Fact Verification

3 benchmarks · 129 papers with code

Keyword Extraction

3 benchmarks · 33 papers with code

Intent Discovery

3 benchmarks · 20 papers with code

Dialogue Rewriting

3 benchmarks · 4 papers with code

Attribute Mining

3 benchmarks · 3 papers with code

POS Tagging

2 benchmarks · 136 papers with code

Legal Reasoning

2 benchmarks · 30 papers with code

Graph-to-Sequence

2 benchmarks · 29 papers with code

Cloze Test

2 benchmarks · 28 papers with code

Rumour Detection

2 benchmarks · 23 papers with code

Subjectivity Analysis

2 benchmarks · 21 papers with code

Passage Re-Ranking

2 benchmarks · 19 papers with code

Meeting Summarization

2 benchmarks · 18 papers with code

Semantic entity labeling

2 benchmarks · 12 papers with code

Nested Mention Recognition

2 benchmarks · 10 papers with code

Arabic Text Diacritization

2 benchmarks · 7 papers with code

Thai Word Segmentation

2 benchmarks · 7 papers with code

Reading Order Detection

2 benchmarks · 4 papers with code

Crowdsourced Text Aggregation

2 benchmarks · 1 paper with code

Negation and Speculation Cue Detection

2 benchmarks · 1 paper with code

Phrase Ranking

2 benchmarks · 1 paper with code

Phrase Tagging

2 benchmarks · 1 paper with code

Memorization

1 benchmark · 438 papers with code

GSM8K

1 benchmark · 209 papers with code

Relational Reasoning

1 benchmark · 179 papers with code

Word Similarity

1 benchmark · 117 papers with code

Passage Ranking

1 benchmark · 35 papers with code

Knowledge Base Population

1 benchmark · 32 papers with code

Automated Essay Scoring

1 benchmark · 27 papers with code

Semantic Retrieval

1 benchmark · 26 papers with code

Sentence Ordering

1 benchmark · 21 papers with code

Humor Detection

1 benchmark · 20 papers with code

Table-based Fact Verification

1 benchmark · 19 papers with code

Code Repair

1 benchmark · 14 papers with code

Dialog Act Classification

1 benchmark · 11 papers with code

Probing Language Models

1 benchmark · 11 papers with code

CCG Supertagging

1 benchmark · 8 papers with code

Fact Selection

1 benchmark · 7 papers with code

Multi-agent Integration

1 benchmark · 7 papers with code

answerability prediction

1 benchmark · 6 papers with code

Aspect Category Polarity

1 benchmark · 5 papers with code

Chinese Spell Checking

1 benchmark · 4 papers with code

Stereotypical Bias Analysis

1 benchmark · 4 papers with code

Zero-shot Sentiment Classification

1 benchmark · 4 papers with code

Action Parsing

1 benchmark · 3 papers with code

Domain Labelling

1 benchmark · 3 papers with code

Japanese Word Segmentation

1 benchmark · 3 papers with code

Memex Question Answering

1 benchmark · 3 papers with code

Polyphone disambiguation

1 benchmark · 3 papers with code

Twitter Event Detection

1 benchmark · 3 papers with code

AMR Graph Similarity

1 benchmark · 2 papers with code

Binary Condescension Detection

1 benchmark · 2 papers with code

Conversational Web Navigation

1 benchmark · 2 papers with code

Croatian Text Diacritization

1 benchmark · 2 papers with code

Czech Text Diacritization

1 benchmark · 2 papers with code

Description-guided molecule generation

1 benchmark · 2 papers with code

French Text Diacritization

1 benchmark · 2 papers with code

Hungarian Text Diacritization

1 benchmark · 2 papers with code

Irish Text Diacritization

1 benchmark · 2 papers with code

Latvian Text Diacritization

1 benchmark · 2 papers with code

Morpheme Segmentaiton

1 benchmark · 2 papers with code

Multi-label Condescension Detection

1 benchmark · 2 papers with code

Personality Recognition in Conversation

1 benchmark · 2 papers with code

Role-filler Entity Extraction

1 benchmark · 2 papers with code

Romanian Text Diacritization

1 benchmark · 2 papers with code

Slovak Text Diacritization

1 benchmark · 2 papers with code

Spanish Text Diacritization

1 benchmark · 2 papers with code

Turkish Text Diacritization

1 benchmark · 2 papers with code

Vietnamese Text Diacritization

1 benchmark · 2 papers with code

Chemical Indexing

1 benchmark · 1 paper with code

Clinical Assertion Status Detection

1 benchmark · 1 paper with code

Commonsense Reasoning for RL

1 benchmark · 1 paper with code

GermEval2024 Shared Task 1 Subtask 1

1 benchmark · 1 paper with code

GermEval2024 Shared Task 1 Subtask 2

1 benchmark · 1 paper with code

Math Information Retrieval

1 benchmark · 1 paper with code

Multimodal Text Prediction

1 benchmark · 1 paper with code

Poem meters classification

1 benchmark · 1 paper with code

Query Wellformedness

1 benchmark · 1 paper with code

Question-Answer categorization

1 benchmark · 1 paper with code

TinyQA Benchmark++

1 benchmark · 1 paper with code

Counterspeech Detection

1 benchmark · 0 papers with code

In-Context Learning

0 benchmarks · 998 papers with code

Retrieval-augmented Generation

0 benchmarks · 777 papers with code

Mamba

0 benchmarks · 552 papers with code

Specificity

0 benchmarks · 504 papers with code

Drug Design

0 benchmarks · 180 papers with code

Text Matching

0 benchmarks · 161 papers with code

Safety Alignment

0 benchmarks · 134 papers with code

Self-Learning

0 benchmarks · 125 papers with code

molecular representation

0 benchmarks · 87 papers with code

Novelty Detection

0 benchmarks · 84 papers with code

Conversational Search

0 benchmarks · 75 papers with code

Morphological Analysis

0 benchmarks · 73 papers with code

Lemmatization

0 benchmarks · 68 papers with code

Abusive Language

0 benchmarks · 51 papers with code

Protein Folding

0 benchmarks · 51 papers with code

Deep Attention

0 benchmarks · 45 papers with code

Multilingual NLP

0 benchmarks · 42 papers with code

Word Translation

0 benchmarks · 40 papers with code

Morphological Inflection

0 benchmarks · 39 papers with code

text annotation

0 benchmarks · 39 papers with code

Authorship Verification

0 benchmarks · 32 papers with code

Cross-Lingual Word Embeddings

0 benchmarks · 32 papers with code

nlg evaluation

0 benchmarks · 31 papers with code

Hallucination Evaluation

0 benchmarks · 29 papers with code

Text Normalization

0 benchmarks · 27 papers with code

X-ray Classification

0 benchmarks · 27 papers with code

Morphological Tagging

0 benchmarks · 26 papers with code

Comment Generation

0 benchmarks · 23 papers with code

Entity Extraction using GAN

0 benchmarks · 22 papers with code

Review Generation

0 benchmarks · 22 papers with code

Sentence-Pair Classification

0 benchmarks · 22 papers with code

Semantic Composition

0 benchmarks · 21 papers with code

Lexical Simplification

0 benchmarks · 20 papers with code

Punctuation Restoration

0 benchmarks · 19 papers with code

Script Generation

0 benchmarks · 18 papers with code

Text Compression

0 benchmarks · 18 papers with code

Reverse Dictionary

0 benchmarks · 17 papers with code

Diachronic Word Embeddings

0 benchmarks · 14 papers with code

Decipherment

0 benchmarks · 12 papers with code

Pretrained Multilingual Language Models

0 benchmarks · 12 papers with code

Linguistic steganography

0 benchmarks · 11 papers with code

Gender Bias Detection

0 benchmarks · 10 papers with code

Human Agent Collaboration

0 benchmarks · 10 papers with code

Clickbait Detection

0 benchmarks · 9 papers with code

Complex Word Identification

0 benchmarks · 8 papers with code

Sign Language Production

0 benchmarks · 8 papers with code

Text Anonymization

0 benchmarks · 8 papers with code

Toponym Resolution

0 benchmarks · 8 papers with code

Vietnamese Datasets

0 benchmarks · 8 papers with code

Commonsense Causal Reasoning

0 benchmarks · 7 papers with code

Vietnamese Hate Speech Detection

0 benchmarks · 7 papers with code

Abstract Argumentation

0 benchmarks · 6 papers with code

Aggression Identification

0 benchmarks · 6 papers with code

Suggestion mining

0 benchmarks · 6 papers with code

Vietnamese Word Segmentation

0 benchmarks · 6 papers with code

Morphological Disambiguation

0 benchmarks · 5 papers with code

Relevance Detection

0 benchmarks · 5 papers with code

Simultaneous Speech-to-Speech Translation

0 benchmarks · 5 papers with code

Text Attribute Transfer

0 benchmarks · 5 papers with code

AI and Safety

0 benchmarks · 4 papers with code

Music Genre Transfer

0 benchmarks · 4 papers with code

WNLI

0 benchmarks · 4 papers with code

Automated Writing Evaluation

0 benchmarks · 3 papers with code

Cognate Prediction

0 benchmarks · 3 papers with code

Social Media Mental Health Detection

0 benchmarks · 3 papers with code

Text-Variation

0 benchmarks · 3 papers with code

Vietnamese Language Models

0 benchmarks · 3 papers with code

Zero-Shot Machine Translation

0 benchmarks · 3 papers with code

ArabicMMLU

0 benchmarks · 2 papers with code

Author Attribution

0 benchmarks · 2 papers with code

Context Query Reformulation

0 benchmarks · 2 papers with code

Definition Modelling

0 benchmarks · 2 papers with code

Hate Span Identification

0 benchmarks · 2 papers with code

Job Prediction

0 benchmarks · 2 papers with code

Misogynistic Aggression Identification

0 benchmarks · 2 papers with code

News Annotation

0 benchmarks · 2 papers with code

Open Relation Modeling

0 benchmarks · 2 papers with code

Record linking

0 benchmarks · 2 papers with code

Self-Evolving AI

0 benchmarks · 2 papers with code

Syntax Representation

0 benchmarks · 2 papers with code

Text-to-video search

0 benchmarks · 2 papers with code

Turning Point Identification

0 benchmarks · 2 papers with code

Vietnamese Scene Text

0 benchmarks · 2 papers with code

Vietnamese Speech Recognition

0 benchmarks · 2 papers with code

Asynchronous Group Communication

0 benchmarks · 1 paper with code

Coding Problem Tagging

0 benchmarks · 1 paper with code

Collaborative Plan Acquisition

0 benchmarks · 1 paper with code

Cross-lingual Text-to-Image Generation

0 benchmarks · 1 paper with code

Emergent communications on relations

0 benchmarks · 1 paper with code

Emotion Detection and Trigger Summarization

0 benchmarks · 1 paper with code

Extractive Tags Summarization

0 benchmarks · 1 paper with code

Hate Intensity Prediction

0 benchmarks · 1 paper with code

incongruity detection

0 benchmarks · 1 paper with code

Joint NER and Classification

0 benchmarks · 1 paper with code

Meme Captioning

0 benchmarks · 1 paper with code

Multi-Dialect Vietnamese

0 benchmarks · 1 paper with code

multi-word expression embedding

0 benchmarks · 1 paper with code

multi-word expression sememe prediction

0 benchmarks · 1 paper with code

Negation and Speculation Scope resolution

0 benchmarks · 1 paper with code

Only Connect Walls Dataset Task 2 (Connections)

0 benchmarks · 1 paper with code

Overlapping Mention Recognition

0 benchmarks · 1 paper with code

Philosophical Reflection

0 benchmarks · 1 paper with code

Phrase Vector Embedding

0 benchmarks · 1 paper with code

Readability optimization

0 benchmarks · 1 paper with code

Reliable Intelligence Identification

0 benchmarks · 1 paper with code

Semi-Supervised Text Regression

0 benchmarks · 1 paper with code

Text Effects Transfer

0 benchmarks · 1 paper with code

Text-to-GQL

0 benchmarks · 1 paper with code

Vietnamese Fact Checking

0 benchmarks · 1 paper with code

Vietnamese Lexical Normalization

0 benchmarks · 1 paper with code

Vietnamese Natural Language Understanding

0 benchmarks · 1 paper with code

Web Page Tagging

0 benchmarks · 1 paper with code

When should a hot water tank be replaced?

0 benchmarks · 1 paper with code

ARQMath2

0 benchmarks · 0 papers with code

Automatic Writing

0 benchmarks · 0 papers with code

Complaint Comment Classification

0 benchmarks · 0 papers with code

Face Selection

0 benchmarks · 0 papers with code

Job classification

0 benchmarks · 0 papers with code

Multi-lingual Text-to-Image Generation

0 benchmarks · 0 papers with code

Multlingual Neural Machine Translation

0 benchmarks · 0 papers with code

Question to Declarative Sentence

0 benchmarks · 0 papers with code

Vietnamese Parsing

0 benchmarks · 0 papers with code

3 tasks in Natural Language Processing are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.

Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.