Home › Datasets › task › Decision Making

Decision Making datasets

archive 2025-07-28

40 datasets carry the task tag "Decision Making" (the task itself: Decision Making), ordered by the archive's paper count. Page 1 of 1: 40 shown of 40. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Decision Making datasets 1–40 of 40

D4RL is a collection of environments for offline reinforcement learning.
538 papers · 2 benchmarks
Charades-STA is a new dataset built on top of Charades by adding sentence temporal annotations.
236 papers · 4 benchmarks
FairFace is a face image dataset which is race balanced.
208 papers · 1 benchmark
ViZDoom is an AI research platform based on the classical First Person Shooter game Doom.
156 papers · 3 benchmarks
LogiQA consists of 8,678 QA instances, covering multiple types of deductive reasoning.
127 papers · 1 benchmark
TORCS (The Open Racing Car Simulator)
TORCS (The Open Racing Car Simulator) is a driving simulator.
96 papers · 0 benchmarks
A large dataset of musculoskeletal radiographs containing 40,561 images from 14,863 studies, where each study is manually labeled by radiologists as either normal or abnormal.
43 papers · 0 benchmarks
ShARC (Shaping Answers with Rules through Conversation)
ShARC is a Conversational Question Answering dataset focussing on question answering from texts containing rules.
43 papers · 0 benchmarks
PearRead is a dataset of scientific peer reviews.
42 papers · 0 benchmarks
CoAID include diverse COVID-19 healthcare misinformation, including fake news on websites and social platforms, along with users' social engagement about such news.
41 papers · 0 benchmarks
GazeFollow is a large-scale dataset annotated with the location of where people in images are looking.
38 papers · 1 benchmark
An open database for sharing robotic experience, which provides an initial pool of 15 million video frames, from 7 different robot platforms, and study how it can be used to learn generalizable models for vision-based robotic manipulation.
28 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
ROAD (ROAD: The ROad event Awareness Dataset for Autonomous Driving)
ROAD is designed to test an autonomous vehicle's ability to detect road events, defined as triplets composed by an active agent, the action(s) it performs and the corresponding scene locations.
27 papers · 0 benchmarks
The SentiCap dataset contains several thousand images with captions with positive and negative sentiments.
26 papers · 0 benchmarks
ROPES (Reasoning Over Paragraph Effects in Situations)
ROPES is a QA dataset which tests a system's ability to apply knowledge from a passage of text to a new situation.
24 papers · 0 benchmarks
Obstacle Tower is a high fidelity, 3D, 3rd person, procedurally generated environment for reinforcement learning.
20 papers · 6 benchmarks
A benchmark which bridges the gap between freely available, documented, and motivated artificial benchmarks and properties of real industrial problems.
13 papers · 0 benchmarks
The TrajNet Challenge represents a large multi-scenario forecasting benchmark.
12 papers · 2 benchmarks
CHALET (Cornell House Agent Learning Environment)
CHALET is a 3D house simulator with support for navigation and manipulation.
10 papers · 0 benchmarks
Atari-HEAD is a dataset of human actions and eye movements recorded while playing Atari videos games.
9 papers · 0 benchmarks
BLVD is a large scale 5D semantics dataset collected by the Visual Cognitive Computing and Intelligent Vehicles Lab.
9 papers · 0 benchmarks
UKP (UKP Argument Annotated Essays)
The UKP Argument Annotated Essays corpus consists of argument annotated persuasive essays including annotations of argument components and argumentative relations.
8 papers · 0 benchmarks
NASA C-MAPSS (Turbofan Engine Degradation Simulation Data Set)
Engine degradation simulation was carried out using C-MAPSS.
7 papers · 2 benchmarks
OMICS (Open Mind Indoor Common Sense)
OMICS is an extensive collection of knowledge for indoor service robots gathered from internet users.
6 papers · 0 benchmarks
Dem@Care is providing the following datasets, which are collected during lab and home experiments.
5 papers · 0 benchmarks
The Flick Cropping Dataset consists of high quality cropping and pairwise ranking annotations used to evaluate the performance of automatic image cropping approaches.
5 papers · 0 benchmarks
GolfDB is a high-quality video dataset created for general recognition applications in the sport of golf, and specifically for the task of golf swing sequencing.
5 papers · 0 benchmarks
CUHK Image Cropping is a dataset for image cropping.
3 papers · 0 benchmarks
The ability to jointly understand the geometry of objects and plan actions for manipulating them is crucial for intelligent agents.
3 papers · 1 benchmark
Covid-HeRA is a dataset for health risk assessment and severity-informed decision making in the presence of COVID19 misinformation.
2 papers · 0 benchmarks
This dataset consists of 5808 dialogues, based on 2236 unique scenarios.
2 papers · 0 benchmarks
PICO is a framework to formulate a well-defined focused clinical question.
2 papers · 0 benchmarks
A View From Somewhere (AVFS)—a dataset of 638,180 face similarity judgments over 4,921 faces.
1 paper · 0 benchmarks
BeNYfits (New York City Public Benefits Eligibility Dialog Agent Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Car_Price_Prediction (Second_Hand-Car_Price_Prediction)
In this dataset we added [Company Name, Car Model, Car Type, Fuel Type, Transmission, Engine (cc), Mileage, Kmsdriven, Buyers, Horsepower (kw), Year Price (Lakhs)]
1 paper · 1 benchmark
A large-scale comprehensive collection of dashcam videos collected by vehicles on DiDi's platform.
1 paper · 0 benchmarks
This is a real-world industrial benchmark dataset from a major medical device manufacturer for the prediction of customer escalations.
1 paper · 0 benchmarks
The Spaceship dataset is a dataset for evaluating agents’ ability to learn to solve a class of physics-based tasks.
1 paper · 0 benchmarks
pursuitMW (Multi-agent pursuit in matrix world)
Multi-agent pursuit in matrix world (pursuitMW) is a partially observable Markov game (POMG) between a swarm of pursuers and a swarm of evaders.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.