Papers › ChatQA: Surpassing GPT-4 on Conversational QA and RAG

ChatQA: Surpassing GPT-4 on Conversational QA and RAG

18 Jan 2024arXiv:2401.10225archive 2025-07-28

Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, Bryan Catanzaro

In this work, we introduce ChatQA, a suite of models that outperform GPT-4 on retrieval-augmented generation (RAG) and conversational question answering (QA). To enhance generation, we propose a two-stage instruction tuning method that significantly boosts the performance of RAG. For effective retrieval, we introduce a dense retriever optimized for conversational QA, which yields results comparable to the alternative state-of-the-art query rewriting models, while substantially reducing deployment costs. We also present the ChatRAG Bench, which encompasses ten datasets covering comprehensive evaluations on RAG, table-related QA, arithmetic calculations, and scenarios involving unanswerable questions. Our ChatQA-1.0-70B (score: 54.14), built on Llama2, a weaker foundation model than GPT-4, can slightly outperform GPT-4-0613 (score: 53.90) and GPT-4-Turbo-2024-04-09 (score: 54.03) on the ChatRAG Bench, without relying on any synthetic data from OpenAI GPT models. Notably, the Llama3-ChatQA-1.5-70B model surpasses the accuracy of GPT-4-Turbo-2024-04-09, achieving a 4.4% improvement. To advance research in this field, we open-sourced the model weights, instruction tuning data, ChatRAG Bench, and retriever for the community: https://chatqa-project.github.io/.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Conversational Question AnsweringQuestion AnsweringRAGRetrievalRetrieval-augmented Generation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Question Answering Natural Questions ChatQA-1.5-llama3-70b (Zero-Shot, KILT) EM 47.0 #13 of 47 Archive leaderboard report
Question Answering Natural Questions ChatQA-1.5-llama3-8b (Zero-Shot, KILT) EM 42.7 #20 of 47 Archive leaderboard report
Question Answering TriviaQA ChatQA-1.5-llama3-70b (Zero-Shot, KILT) EM 85.6 #6 of 56 Archive leaderboard report
Question Answering TriviaQA ChatQA-1.5-llama3-8B (Zero-Shot, KILT) EM 81.0 #13 of 56 Archive leaderboard report
Question Answering TriviaQA ChatQA-1.5-llama3-70b (Zero-Shot, DPR) EM 69.0 #33 of 56 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionAttention DropoutBARTBERTBPECosine AnnealingDense ConnectionsDiscriminative Fine-TuningDropoutGPTGPT-4Label SmoothingLayer NormalizationLinear LayerLinear Warmup With Cosine AnnealingLinear Warmup With Linear DecayMulti-Head AttentionPosition-Wise Feed-Forward LayerRAGResidual ConnectionSoftmaxTransformerWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections