Papers › GPT-4 Technical Report
GPT-4 Technical Report
15 Mar 2023Preprint 2023 3arXiv:2303.08774archive 2025-07-28
OpenAI, :, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny Bogdonoff, Oleg Boiko, Madelaine Boyd, Anna-Luisa Brakman, Greg Brockman, Tim Brooks, Miles Brundage, Kevin Button, Trevor Cai, Rosie Campbell, Andrew Cann, Brittany Carey, Chelsea Carlson, Rory Carmichael, Brooke Chan, Che Chang, Fotis Chantzis, Derek Chen, Sully Chen, Ruby Chen, Jason Chen, Mark Chen, Ben Chess, Chester Cho, Casey Chu, Hyung Won Chung, Dave Cummings, Jeremiah Currier, Yunxing Dai, Cory Decareaux, Thomas Degry, Noah Deutsch, Damien Deville, Arka Dhar, David Dohan, Steve Dowling, Sheila Dunning, Adrien Ecoffet, Atty Eleti, Tyna Eloundou, David Farhi, Liam Fedus, Niko Felix, Simón Posada Fishman, Juston Forte, Isabella Fulford, Leo Gao, Elie Georges, Christian Gibson, Vik Goel, Tarun Gogineni, Gabriel Goh, Rapha Gontijo-Lopes, Jonathan Gordon, Morgan Grafstein, Scott Gray, Ryan Greene, Joshua Gross, Shixiang Shane Gu, Yufei Guo, Chris Hallacy, Jesse Han, Jeff Harris, Yuchen He, Mike Heaton, Johannes Heidecke, Chris Hesse, Alan Hickey, Wade Hickey, Peter Hoeschele, Brandon Houghton, Kenny Hsu, Shengli Hu, Xin Hu, Joost Huizinga, Shantanu Jain, Shawn Jain, Joanne Jang, Angela Jiang, Roger Jiang, Haozhun Jin, Denny Jin, Shino Jomoto, Billie Jonn, Heewoo Jun, Tomer Kaftan, Łukasz Kaiser, Ali Kamali, Ingmar Kanitscheider, Nitish Shirish Keskar, Tabarak Khan, Logan Kilpatrick, Jong Wook Kim, Christina Kim, Yongjik Kim, Jan Hendrik Kirchner, Jamie Kiros, Matt Knight, Daniel Kokotajlo, Łukasz Kondraciuk, Andrew Kondrich, Aris Konstantinidis, Kyle Kosic, Gretchen Krueger, Vishal Kuo, Michael Lampe, Ikai Lan, Teddy Lee, Jan Leike, Jade Leung, Daniel Levy, Chak Ming Li, Rachel Lim, Molly Lin, Stephanie Lin, Mateusz Litwin, Theresa Lopez, Ryan Lowe, Patricia Lue, Anna Makanju, Kim Malfacini, Sam Manning, Todor Markov, Yaniv Markovski, Bianca Martin, Katie Mayer, Andrew Mayne, Bob McGrew, Scott Mayer McKinney, Christine McLeavey, Paul McMillan, Jake McNeil, David Medina, Aalok Mehta, Jacob Menick, Luke Metz, Andrey Mishchenko, Pamela Mishkin, Vinnie Monaco, Evan Morikawa, Daniel Mossing, Tong Mu, Mira Murati, Oleg Murk, David Mély, Ashvin Nair, Reiichiro Nakano, Rajeev Nayak, Arvind Neelakantan, Richard Ngo, Hyeonwoo Noh, Long Ouyang, Cullen O'Keefe, Jakub Pachocki, Alex Paino, Joe Palermo, Ashley Pantuliano, Giambattista Parascandolo, Joel Parish, Emy Parparita, Alex Passos, Mikhail Pavlov, Andrew Peng, Adam Perelman, Filipe de Avila Belbute Peres, Michael Petrov, Henrique Ponde de Oliveira Pinto, Michael, Pokorny, Michelle Pokrass, Vitchyr H. Pong, Tolly Powell, Alethea Power, Boris Power, Elizabeth Proehl, Raul Puri, Alec Radford, Jack Rae, Aditya Ramesh, Cameron Raymond, Francis Real, Kendra Rimbach, Carl Ross, Bob Rotsted, Henri Roussez, Nick Ryder, Mario Saltarelli, Ted Sanders, Shibani Santurkar, Girish Sastry, Heather Schmidt, David Schnurr, John Schulman, Daniel Selsam, Kyla Sheppard, Toki Sherbakov, Jessica Shieh, Sarah Shoker, Pranav Shyam, Szymon Sidor, Eric Sigler, Maddie Simens, Jordan Sitkin, Katarina Slama, Ian Sohl, Benjamin Sokolowsky, Yang song, Natalie Staudacher, Felipe Petroski Such, Natalie Summers, Ilya Sutskever, Jie Tang, Nikolas Tezak, Madeleine B. Thompson, Phil Tillet, Amin Tootoonchian, Elizabeth Tseng, Preston Tuggle, Nick Turley, Jerry Tworek, Juan Felipe Cerón Uribe, Andrea Vallone, Arun Vijayvergiya, Chelsea Voss, Carroll Wainwright, Justin Jay Wang, Alvin Wang, Ben Wang, Jonathan Ward, Jason Wei, CJ Weinmann, Akila Welihinda, Peter Welinder, Jiayi Weng, Lilian Weng, Matt Wiethoff, Dave Willner, Clemens Winter, Samuel Wolrich, Hannah Wong, Lauren Workman, Sherwin Wu, Jeff Wu, Michael Wu, Kai Xiao, Tao Xu, Sarah Yoo, Kevin Yu, Qiming Yuan, Wojciech Zaremba, Rowan Zellers, Chong Zhang, Marvin Zhang, Shengjia Zhao, Tianhao Zheng, Juntang Zhuang, William Zhuk, Barret Zoph
We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs. While less capable than humans in many real-world scenarios, GPT-4 exhibits human-level performance on various professional and academic benchmarks, including passing a simulated bar exam with a score around the top 10% of test takers. GPT-4 is a Transformer-based model pre-trained to predict the next token in a document. The post-training alignment process results in improved performance on measures of factuality and adherence to desired behavior. A core component of this project was developing infrastructure and optimization methods that behave predictably across a wide range of scales. This allowed us to accurately predict some aspects of GPT-4's performance based on models trained with no more than 1/1,000th the compute of GPT-4.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2303.08774")
Code
Syntology Ran 2 of 5 code samples harvested from 2 repositories linked to this paper; 3 have no recorded run. Of those that ran: 2 ran · violated contract.
By repository: community (archive-listed): 4 samples from 2 repositories, 1 ran; 1 identical to code first harvested elsewhere. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
5 samples harvested; 2 ran; 0 honoured the contract we drafted; 3 have no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.
2ran · violated contract
3unverified
Licence: 1 of the 5 samples is pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.
Harvested from 2 repositories linked to this paper, official or community; each sample names its own and says which. Some samples are identical code Syntology first harvested from another repository; for those, this paper's copy is not located and its licence is not recorded. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.
Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.
Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.
parse_score
identical code first harvested elsewhere
ran · violated contract
fingerprinted
licence of this copy not recorded · b748891b4b063b83 · report
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
| Arithmetic Reasoning |
GSM8K |
GPT-3.5 (few-shot, k=5) |
Accuracy |
57.1 |
#121 of 164 |
Archive leaderboard |
report |
| Common Sense Reasoning |
ARC (Challenge) |
GPT-4 (few-shot, k=25) |
Accuracy |
96.4 |
#1 of 54 |
Archive leaderboard |
report |
| Common Sense Reasoning |
ARC (Challenge) |
GPT-3.5 (few-shot, k=25) |
Accuracy |
85.2 |
#14 of 54 |
Archive leaderboard |
report |
| Common Sense Reasoning |
WinoGrande |
GPT-4 (5-shot) |
Accuracy |
87.5 |
#7 of 77 |
Archive leaderboard |
report |
| Common Sense Reasoning |
WinoGrande |
GPT-3.5 (5-shot) |
Accuracy |
81.6 |
#14 of 77 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
#Learning Samples (N) |
16 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
ACC |
42.30 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
BLEU-4 |
45.51 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
CIDEr |
269.68 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
Detection |
7.00 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
METEOR |
35.17 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
ROUGE-L |
52.67 |
#2 of 7 |
Archive leaderboard |
report |
| FS-MEVQA |
SME |
GPT-4-1106-Vision-Preview |
SPICE |
37.67 |
#2 of 7 |
Archive leaderboard |
report |
| Factual Inconsistency Detection in Chart Captioning |
CHOCOLATE-LLM |
GPT-4V |
Kendall's Tau-c |
0.205 |
#1 of 5 |
Archive leaderboard |
report |
| Few-Shot Learning |
MedConceptsQA |
gpt-4-0125-preview |
Accuracy |
61.911 |
#1 of 12 |
Archive leaderboard |
report |
| Legal Reasoning |
LegalBench (Rule-recall) |
GPT-4 |
Balanced Accuracy |
59.2 |
#1 of 1 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
128k |
0.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
12k |
49.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
16k |
44.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
1k |
74.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
2k |
73.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
32k |
16.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
4k |
67.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
64k |
0.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
6k |
59.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-1106 |
8k |
53.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
128k |
0.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
12k |
52.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
16k |
44.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
1k |
73.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
2k |
73.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
32k |
30.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
4k |
65.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
64k |
0.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
6k |
63.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (BestAnswer) |
GPT-4-Turbo-0125 |
8k |
56.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
128k |
6.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
16k |
3.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
2k |
18.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
32k |
6.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
4k |
15.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
64k |
6.0 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-1106 |
8k |
7.5 |
#1 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
128k |
2.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
16k |
5.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
2k |
15.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
32k |
2.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
4k |
16.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
64k |
4.0 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
Ada-LEval (TSort) |
GPT-4-Turbo-0125 |
8k |
8.5 |
#2 of 10 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
1 Image, 2*2 Stitching, Exact Accuracy |
94.6 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
1 Image, 4*4 Stitching, Exact Accuracy |
83 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
1 Image, 8*8 Stitching, Exact Accuracy |
19 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
10 Images, 1*1 Stitching, Exact Accuracy |
97 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
10 Images, 2*2 Stitching, Exact Accuracy |
81.8 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
10 Images, 4*4 Stitching, Exact Accuracy |
26.9 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4o |
10 Images, 8*8 Stitching, Exact Accuracy |
1 |
#1 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
1 Image, 2*2 Stitching, Exact Accuracy |
86.09 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
1 Image, 4*4 Stitching, Exact Accuracy |
54.72 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
1 Image, 8*8 Stitching, Exact Accuracy |
7.3 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
10 Images, 1*1 Stitching, Exact Accuracy |
72.36 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
10 Images, 2*2 Stitching, Exact Accuracy |
34.24 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
10 Images, 4*4 Stitching, Exact Accuracy |
7.58 |
#2 of 12 |
Archive leaderboard |
report |
| Long-Context Understanding |
MMNeedle |
GPT-4V |
10 Images, 8*8 Stitching, Exact Accuracy |
0 |
#2 of 12 |
Archive leaderboard |
report |
| Multi-task Language Understanding |
MML |
GPT-3.5 Turbo |
Average (%) |
70.0 |
#12 of 44 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
6-DoF |
- |
#4 of 5 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
pos-level0 |
39.1 |
#4 of 5 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
pos-level1 |
46.8 |
#4 of 5 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
rot-level0 |
9.1 |
#4 of 5 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
rot-level1 |
6.9 |
#4 of 5 |
Archive leaderboard |
report |
| Object Rearrangement |
Open6DOR V2 |
GPT-4V |
rot-level2 |
11.7 |
#4 of 5 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
Wasserstein Distance (WD) |
72.9 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
# Correct Groups |
269 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
# Solved Walls |
7 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
Adjusted Mutual Information (AMI) |
32.8 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
Adjusted Rand Index (ARI) |
29.1 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (5-shot) |
Fowlkes Mallows Score (FMS) |
43.4 |
#1 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
Wasserstein Distance (WD) |
73.4 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
# Correct Groups |
262 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
# Solved Walls |
4 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
Adjusted Mutual Information (AMI) |
33.5 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
Adjusted Rand Index (ARI) |
29.7 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (1-shot) |
Fowlkes Mallows Score (FMS) |
43.7 |
#2 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
Wasserstein Distance (WD) |
73.6 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
# Correct Groups |
249 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
# Solved Walls |
3 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
Adjusted Mutual Information (AMI) |
32.3 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
Adjusted Rand Index (ARI) |
28.5 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (100-shot) |
Fowlkes Mallows Score (FMS) |
42.8 |
#3 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
Wasserstein Distance (WD) |
73.7 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
# Correct Groups |
272 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
# Solved Walls |
5 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
Adjusted Mutual Information (AMI) |
33.6 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
Adjusted Rand Index (ARI) |
29.9 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (3-shot) |
Fowlkes Mallows Score (FMS) |
43.9 |
#4 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
Wasserstein Distance (WD) |
75.8 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
# Correct Groups |
239 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
# Solved Walls |
6 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
Adjusted Mutual Information (AMI) |
30.7 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
Adjusted Rand Index (ARI) |
27.2 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-4 (0-shot) |
Fowlkes Mallows Score (FMS) |
41.5 |
#5 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
Wasserstein Distance (WD) |
80.6 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
# Correct Groups |
149 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
# Solved Walls |
2 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
Adjusted Mutual Information (AMI) |
25.4 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
Adjusted Rand Index (ARI) |
22.0 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (5-shot) |
Fowlkes Mallows Score (FMS) |
37.3 |
#6 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
Wasserstein Distance (WD) |
80.9 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
# Correct Groups |
140 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
# Solved Walls |
0 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
Adjusted Mutual Information (AMI) |
24.7 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
Adjusted Rand Index (ARI) |
21.3 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (3-shot) |
Fowlkes Mallows Score (FMS) |
36.8 |
#7 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
Wasserstein Distance (WD) |
81.2 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
# Correct Groups |
137 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
# Solved Walls |
2 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
Adjusted Mutual Information (AMI) |
24.0 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
Adjusted Rand Index (ARI) |
20.4 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (10-shot) |
Fowlkes Mallows Score (FMS) |
36.1 |
#8 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
Wasserstein Distance (WD) |
82.3 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
# Correct Groups |
123 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
# Solved Walls |
0 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
Adjusted Mutual Information (AMI) |
21.2 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
Adjusted Rand Index (ARI) |
18.2 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (1-shot) |
Fowlkes Mallows Score (FMS) |
34.4 |
#9 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
Wasserstein Distance (WD) |
82.5 |
#10 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
# Correct Groups |
114 |
#10 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
# Solved Walls |
0 |
#10 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
Adjusted Mutual Information (AMI) |
21.6 |
#10 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
Adjusted Rand Index (ARI) |
18.4 |
#10 of 22 |
Archive leaderboard |
report |
| Only Connect Walls Dataset Task 1 (Grouping) |
OCW |
GPT-3.5-turbo (0-shot) |
Fowlkes Mallows Score (FMS) |
34.0 |
#10 of 22 |
Archive leaderboard |
report |
| Question Answering |
DROP Test |
GPT-4 (few-shot, k=3) |
F1 |
80.9 |
#6 of 16 |
Archive leaderboard |
report |
| Question Answering |
DROP Test |
GPT 3.5 (few-shot, k=3) |
F1 |
64.1 |
#11 of 16 |
Archive leaderboard |
report |
| Question Answering |
PeerQA |
GPT-4o-2024-08-06-128k |
AlignScore |
0.1224 |
#1 of 6 |
Archive leaderboard |
report |
| Question Answering |
PeerQA |
GPT-4o-2024-08-06-128k |
Prometheus-2 Answer Correctness |
3.4612 |
#1 of 6 |
Archive leaderboard |
report |
| Question Answering |
PeerQA |
GPT-4o-2024-08-06-128k |
Rouge-L |
0.2266 |
#1 of 6 |
Archive leaderboard |
report |
| Question Answering |
TIQ |
Gpt-4 |
P@1 |
28.6 |
#4 of 9 |
Archive leaderboard |
report |
| Question Answering |
TriviaQA |
GPT-4-0613 (Zero-shot) |
EM |
84.8 |
#8 of 56 |
Archive leaderboard |
report |
| Question Answering |
TruthfulQA |
GPT-4 (RLHF) |
MC1 |
0.59 |
#1 of 33 |
Archive leaderboard |
report |
| Sentence Completion |
HellaSwag |
GPT-4 (10-shot) |
Accuracy |
95.3 |
#4 of 89 |
Archive leaderboard |
report |
| Sentence Completion |
HellaSwag |
GPT-3.5 (10-shot) |
Accuracy |
85.5 |
#23 of 89 |
Archive leaderboard |
report |
| Spatial Reasoning |
EmbSpatial-Bench |
GPT-4V |
Generation |
36.07 |
#3 of 5 |
Archive leaderboard |
report |
| Visual Question Answering |
BenchLMM |
GPT-4V |
GPT-3.5 score |
58.37 |
#1 of 10 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet |
GPT-4o (gpt-4o-2024-05-13) |
GPT-4 score |
69.3±0.1 |
#12 of 231 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet |
gpt-4o-mini-2024-07-18 |
GPT-4 score |
68.6±0.1 |
#14 of 231 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet |
GPT-4V |
GPT-4 score |
67.7±0.3 |
#15 of 231 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet |
GPT-4V-Turbo-detail:high |
GPT-4 score |
67.6±0.1 |
#16 of 231 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet |
GPT-4V-Turbo-detail:low |
GPT-4 score |
60.2±0.3 |
#39 of 231 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet v2 |
GPT-4o (gpt-4o-2024-11-20) |
GPT-4 score |
72.1±0.2 |
#2 of 24 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet v2 |
GPT-4o (gpt-4o-2024-05-13) |
GPT-4 score |
71.0±0.2 |
#4 of 24 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet v2 |
gpt-4o-mini-2024-07-18 |
GPT-4 score |
66.8±0.3 |
#8 of 24 |
Archive leaderboard |
report |
| Visual Question Answering |
MM-Vet v2 |
GPT-4 Turbo (gpt-4-0125-preview) |
GPT-4 score |
66.3±0.2 |
#9 of 24 |
Archive leaderboard |
report |
| Visual Question Answering |
ViP-Bench |
GPT-4V-turbo-detail:high (Visual Prompt) |
GPT-4 score (bbox) |
60.7 |
#1 of 13 |
Archive leaderboard |
report |
| Visual Question Answering |
ViP-Bench |
GPT-4V-turbo-detail:high (Visual Prompt) |
GPT-4 score (human) |
59.9 |
#1 of 13 |
Archive leaderboard |
report |
| Visual Question Answering |
ViP-Bench |
GPT-4V-turbo-detail:low (Visual Prompt) |
GPT-4 score (bbox) |
52.8 |
#2 of 13 |
Archive leaderboard |
report |
| Visual Question Answering |
ViP-Bench |
GPT-4V-turbo-detail:low (Visual Prompt) |
GPT-4 score (human) |
51.4 |
#2 of 13 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
CORE-MM |
GPT-4V |
Abductive |
77.88 |
#1 of 1 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
CORE-MM |
GPT-4V |
Analogical |
69.86 |
#1 of 1 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
CORE-MM |
GPT-4V |
Deductive |
74.86 |
#1 of 1 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
CORE-MM |
GPT-4V |
Overall score |
74.44 |
#1 of 1 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
CORE-MM |
GPT-4V |
Params |
- |
#1 of 1 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
InfiMM-Eval |
GPT-4V |
Abductive |
77.88 |
#1 of 14 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
InfiMM-Eval |
GPT-4V |
Analogical |
69.86 |
#1 of 14 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
InfiMM-Eval |
GPT-4V |
Deductive |
74.86 |
#1 of 14 |
Archive leaderboard |
report |
| Visual Question Answering (VQA) |
InfiMM-Eval |
GPT-4V |
Overall score |
74.44 |
#1 of 14 |
Archive leaderboard |
report |
| Zero-Shot Learning |
MedConceptsQA |
gpt-4-0125-preview |
Accuracy |
52.489 |
#1 of 13 |
Archive leaderboard |
report |
| answerability prediction |
PeerQA |
GPT-4o-2024-08-06 |
Macro F1 |
0.3087 |
#5 of 6 |
Archive leaderboard |
report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Introduced by this paper: GPT-4
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections