Home › Datasets › task › AI Agent

AI Agent datasets

archive 2025-07-28

8 datasets carry the task tag "AI Agent" (the task itself: AI Agent), ordered by the archive's paper count. Page 1 of 1: 8 shown of 8. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

AI Agent datasets 1–8 of 8

A multimodal agent benchmark on professional data science and engineering.
7 papers · 0 benchmarks
DEVAI is a benchmark of 55 realistic AI development tasks.
4 papers · 0 benchmarks
BeNYfits (New York City Public Benefits Eligibility Dialog Agent Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
CAGUI (Chinese Android GUI Benchmark)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
To evaluate our proposed strategy of asynchronous communication for LLMs, we run games of Mafia with human players, incorporating an LLM-based agent as an additional player, within an asynchronous chat environment.
1 paper · 0 benchmarks
MMTB (Multi-Mission Tool Bench)
Our test data has undergone five rounds of manual inspection and correction by five senior algorithm researcher with years of experience in NLP, CV, and LLM, taking about one month in total.
1 paper · 0 benchmarks
An evaluation dataset for planning with LLM agents
1 paper · 0 benchmarks
diaforge-utc-r-0725 (DiaFORGE UTC: Unified Tool-Calling Conversations Dataset)
Dataset for our paper Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky which includes 5000 enterprise tools and the corresponding dialogues generated using DiaFORGE UTC data engine.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.