Datasets › Yelp

Yelp

archive 2025-07-28

The Yelp Dataset is a valuable resource for academic research, teaching, and learning. It provides a rich collection of real-world data related to businesses, reviews, and user interactions. Here are the key details about the Yelp Dataset: Reviews: A whopping 6,990,280 reviews from users. Businesses: Information on 150,346 businesses. Pictures: A collection of 200,100 pictures. Metropolitan Areas: Data from 11 metropolitan areas. Tips: Over 908,915 tips provided by 1,987,897 users. Business Attributes: Details like hours, parking availability, and ambiance for more than 1.2 million businesses. Aggregated Check-ins: Historical check-in data for each of the 131,930 businesses.

Benchmarks archive 2025-07-28

All 15 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Sentiment Analysis Yelp Binary classification XLNet Error 1.37 XLNet: Generalized Autoregressive Pretraining for... huggingface/transformers +26 20 Compare
Sentiment Analysis Yelp Fine-grained classification XLNet Error 27.05 XLNet: Generalized Autoregressive Pretraining for... huggingface/transformers +26 17 Compare
Link Prediction Yelp PEAGAT HR@10 0.9128 Metapath- and Entity-aware Graph Neural Network for... ecml-peagnn/PEAGNN 9 Compare
Text Style Transfer Yelp Review Dataset (Small) SAE+Discriminator G-Score (BLEU, Accuracy) 74.56 Style Transfer for Texts: Retrain, Report Errors,... VAShibaev/text_style_transfer 8 Compare
Text Classification Yelp-5 HAHNN (CNN) Accuracy 73.28% Hierarchical Attentional Hybrid Neural Networks for... luisfredgs/cnn-hierarchical-network-for-document-classification +1 7 Compare
Text Classification Yelp-2 XLNet Accuracy 98.63% XLNet: Generalized Autoregressive Pretraining for... huggingface/transformers +26 5 Compare
Multibehavior Recommendation Yelp HMAR HR@10 0.9015 HMAR: Hierarchical Masked Attention for Multi-Behaviour... shereen-elsayed/hmar 4 Compare
Unsupervised Opinion Summarization Yelp BiMeanVAE - Coop ROUGE-1 35.37 Convex Aggregation for Opinion Summarization megagonlabs/coop 4 Compare
Recommendation Systems Yelp DGRec NDCG 0.1427 Session-based Social Recommendation via Dynamic Graph... DeepGraphLearning/RecommenderSystems +1 2 Compare
Text Style Transfer Yelp Review Dataset (Large) SentiInc G-Score (BLEU, Accuracy) 59.17 SentiInc: Incorporating Sentiment Information into... — 2 Compare
Aspect Extraction YASO - YELP RACL - Laptops F1 23 YASO: A Targeted Sentiment Analysis Evaluation Dataset... IBM/yaso-tsa +1 1 Compare
Document Classification Yelp-14 KD-LSTMreg Accuracy 69.4 DocBERT: BERT for Document Classification castorini/hedwig +2 1 Compare
Paraphrase Identification Yelp SplitEE-S Accuracy 76.7 SplitEE: Early Exit in Deep Neural Networks with Split Computing Div290/SplitEE 1 Compare
Sequential Recommendation Yelp BSARec HR@5 0.0275 An Attentive Inductive Bias for Sequential... yehjin-shin/bsarec +1 1 Compare
Text Classification yelp_polarity no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 49 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 86. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Multi-Behavioral Sequential Recommendation 1 1 8 Oct 2024 not harvested
HMAR: Hierarchical Masked Attention for Multi-Behaviour Recommendation 1 3 29 Apr 2024 not harvested
An Attentive Inductive Bias for Sequential Recommendation beyond the Self-Attention 2 1 16 Dec 2023 ran 7 of 7 samples (0 unverified; 7 pointer-only for licence)
SplitEE: Early Exit in Deep Neural Networks with Split Computing 1 1 17 Sep 2023 not harvested
Composable Text Controls in Latent Space with ODEs 1 1 1 Aug 2022 not harvested
A Novel Estimator of Mutual Information for Learning to Disentangle Textual Representations 0 1 6 May 2021 not harvested
Convex Aggregation for Opinion Summarization 1 4 3 Apr 2021 not harvested
YASO: A Targeted Sentiment Analysis Evaluation Dataset for Open-Domain Reviews 2 1 29 Dec 2020 not harvested
Metapath- and Entity-aware Graph Neural Network for Recommendation 1 1 22 Oct 2020 not harvested
Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training Approach 1 1 15 Oct 2020 ran 0 of 5 samples (5 unverified)
Big Bird: Transformers for Longer Sequences 14 1 28 Jul 2020 ran 10 of 15 samples (5 unverified; 11 pointer-only for licence)
Neural Collaborative Filtering vs. Matrix Factorization Revisited 4 1 19 May 2020 not harvested
SentiInc: Incorporating Sentiment Information into Sentiment Transfer Without Parallel Data 0 2 8 Apr 2020 not harvested
Heavy-tailed Representations, Text Polarity Classification & Data Augmentation 0 1 25 Mar 2020 not harvested
Sampling Bias in Deep Active Classification: An Empirical Study 2 2 20 Sep 2019 ran 0 of 1 samples (1 unverified)
Style Transfer for Texts: Retrain, Report Errors, Compare with Rewrites 1 1 19 Aug 2019 ran 0 of 1 samples (1 unverified)
XLNet: Generalized Autoregressive Pretraining for Language Understanding 27 4 19 Jun 2019 ran 10 of 24 samples (14 unverified; 3 pointer-only for licence)
Rethinking Complex Neural Network Architectures for Document Classification 1 1 1 Jun 2019 not harvested
KGAT: Knowledge Graph Attention Network for Recommendation 7 1 20 May 2019 ran 7 of 13 samples (6 unverified; 1 pointer-only for licence)
How to Fine-Tune BERT for Text Classification? 15 6 14 May 2019 ran 6 of 18 samples (12 unverified; 5 pointer-only for licence)
Unsupervised Data Augmentation for Consistency Training 20 6 29 Apr 2019 ran 15 of 52 samples (37 unverified; 9 pointer-only for licence)
DocBERT: BERT for Document Classification 3 1 17 Apr 2019 ran 0 of 14 samples (14 unverified)
Knowledge Graph Convolutional Networks for Recommender Systems 8 1 18 Mar 2019 ran 1 of 5 samples (4 unverified; 1 pointer-only for licence)
Session-based Social Recommendation via Dynamic Graph Attention Networks 2 1 25 Feb 2019 ran 1 of 13 samples (12 unverified)
Learning Topological Representation for Networks via Hierarchical Sampling 1 1 15 Feb 2019 not harvested
Representation Learning for Heterogeneous Information Networks via Embedding Events 1 1 29 Jan 2019 not harvested
Squeezed Very Deep Convolutional Neural Networks for Text Classification 1 2 28 Jan 2019 not harvested
Hierarchical Attentional Hybrid Neural Networks for Document Classification 2 1 20 Jan 2019 not harvested
Learning to Remember More with Less Memorization 1 2 5 Jan 2019 ran 3 of 3 samples (0 unverified)
Compositional Coding Capsule Network with K-Means Routing for Text Classification 1 2 22 Oct 2018 not harvested

The full list of 49 is in the JSON twin.

Dataset loaders archive 2025-07-28

SaeedaAiyub/Task2tfpytorchjax

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Yelp Review Polarity
  • yelp_review_full yelp_review_full
  • yelp_review_full
  • Yelp Review Dataset (Small)
  • Yelp Review Dataset (Large)
  • yelp_polarity
  • Yelp-Fraud
  • Yelp Fine-grained classification
  • Yelp Binary classification
  • Yelp (Amazon en train)
  • Yelp-5
  • Yelp2018
  • Yelp 2014
  • Yelp 2013
  • Yelp-2
  • Yelp15
  • Yelp-14
  • YASO - YELP
  • Yelp

19 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections