Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 27
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 27 of 29: papers 2,601 to 2,700 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Unleashing the Emergent Cognitive Synergy in Large Language Models: A Task-Solving Agent through Multi-Persona Self-Collaboration 11 Jul 2023 · 2 repositories · arXiv:2307.05300
-
ChatGPT for Digital Forensic Investigation: The Good, The Bad, and The Unknown 10 Jul 2023 · 1 repository · arXiv:2307.10195
-
Assessing the efficacy of large language models in generating accurate teacher responses 9 Jul 2023 · 0 repositories · arXiv:2307.04274
-
SVIT: Scaling up Visual Instruction Tuning 9 Jul 2023 · 2 repositories · arXiv:2307.04087
-
Exploring and Characterizing Large Language Models For Embedded System Development and Debugging 7 Jul 2023 · 0 repositories · arXiv:2307.03817
-
Large Language Models as Batteries-Included Zero-Shot ESCO Skills Matchers 7 Jul 2023 · 0 repositories · arXiv:2307.03539
-
Teaching Arithmetic to Small Transformers 7 Jul 2023 · 1 repository · arXiv:2307.03381Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Building Cooperative Embodied Agents Modularly with Large Language Models 5 Jul 2023 · 2 repositories · arXiv:2307.02485
-
Comparative Analysis of GPT-4 and Human Graders in Evaluating Praise Given to Students in Synthetic Dialogues 5 Jul 2023 · 0 repositories · arXiv:2307.02018
-
External Reasoning: Towards Multi-Large-Language-Models Interchangeable Assistance with Human Feedback 5 Jul 2023 · 1 repository · arXiv:2307.12057
-
Hoodwinked: Deception and Cooperation in a Text-Based Game for Language Models 5 Jul 2023 · 1 repository · arXiv:2308.01404
-
Jailbroken: How Does LLM Safety Training Fail? 5 Jul 2023 · 1 repository · arXiv:2307.02483
-
Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning 5 Jul 2023 · 0 repositories · arXiv:2307.02179
-
Evaluating Shutdown Avoidance of Language Models in Textual Scenarios 3 Jul 2023 · 1 repository · arXiv:2307.00787
-
Harnessing LLMs in Curricular Design: Using GPT-4 to Support Authoring of Learning Objectives 30 Jun 2023 · 0 repositories · arXiv:2306.17459
-
Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting 30 Jun 2023 · 0 repositories · arXiv:2306.17563
-
Preference Ranking Optimization for Human Alignment 30 Jun 2023 · 1 repository · arXiv:2306.17492
-
SummQA at MEDIQA-Chat 2023:In-Context Learning with GPT-4 for Medical Summarization 30 Jun 2023 · 1 repository · arXiv:2306.17384
-
A negation detection assessment of GPTs: analysis with the xNot360 dataset 29 Jun 2023 · 0 repositories · arXiv:2306.16638
-
CMATH: Can Your Language Model Pass Chinese Elementary School Math Test? 29 Jun 2023 · 0 repositories · arXiv:2306.16636
-
Generative AI for Programming Education: Benchmarking ChatGPT, GPT-4, and Human Tutors 29 Jun 2023 · 0 repositories · arXiv:2306.17156
-
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding 29 Jun 2023 · 2 repositories · arXiv:2306.17107Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
LyricWhiz: Robust Multilingual Zero-shot Lyrics Transcription by Whispering to ChatGPT 29 Jun 2023 · 1 repository · arXiv:2306.17103
-
UMASS_BioNLP at MEDIQA-Chat 2023: Can LLMs generate high-quality synthetic note-oriented doctor-patient conversations? 29 Jun 2023 · 1 repository · arXiv:2306.16931Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Pareto Optimal Learning for Estimating Large Language Model Errors 28 Jun 2023 · 0 repositories · arXiv:2306.16564
-
Chatlaw: A Multi-Agent Collaborative Legal Assistant with Knowledge Graph Enhanced Mixture-of-Experts Large Language Model 28 Jun 2023 · 1 repository · arXiv:2306.16092
-
Is ChatGPT a Biomedical Expert? -- Exploring the Zero-Shot Performance of Current GPT Models in Biomedical Tasks 28 Jun 2023 · 1 repository · arXiv:2306.16108
-
Leveraging GPT-4 for Food Effect Summarization to Enhance Product-Specific Guidance Development via Iterative Prompting 28 Jun 2023 · 0 repositories · arXiv:2306.16275
-
Taqyim: Evaluating Arabic NLP Tasks Using ChatGPT Models 28 Jun 2023 · 1 repository · arXiv:2306.16322
-
Evaluating GPT-3.5 and GPT-4 on Grammatical Error Correction for Brazilian Portuguese 27 Jun 2023 · 0 repositories · arXiv:2306.15788
-
LeanDojo: Theorem Proving with Retrieval-Augmented Language Models 27 Jun 2023 · 3 repositories · arXiv:2306.15626Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Large Multimodal Models: Notes on CVPR 2023 Tutorial 26 Jun 2023 · 0 repositories · arXiv:2306.14895
-
LM4HPC: Towards Effective Language Model Application in High-Performance Computing 26 Jun 2023 · 0 repositories · arXiv:2306.14979
-
Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise? 24 Jun 2023 · 0 repositories · arXiv:2306.13906
-
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs 22 Jun 2023 · 1 repository · arXiv:2306.13063Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Cross-lingual Cross-temporal Summarization: Dataset, Models, Evaluation 22 Jun 2023 · 1 repository · arXiv:2306.12916
-
Visual Adversarial Examples Jailbreak Aligned Large Language Models 22 Jun 2023 · 1 repository · arXiv:2306.13213
-
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews 21 Jun 2023 · 1 repository · arXiv:2306.12587Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Joint Prompt Optimization of Stacked LLMs using Variational Inference 21 Jun 2023 · 1 repository · arXiv:2306.12509Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 16 unverified (of 19 harvested samples)
-
GPT-Based Models Meet Simulation: How to Efficiently Use Large-Scale Pre-Trained Language Models Across Simulation Tasks 21 Jun 2023 · 0 repositories · arXiv:2306.13679
-
DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models 20 Jun 2023 · 0 repositories · arXiv:2306.11698
-
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts 20 Jun 2023 · 0 repositories · arXiv:2306.11372
-
A GPT-4 Reticular Chemist for Guiding MOF Discovery 20 Jun 2023 · 1 repository · arXiv:2306.14915
-
Harnessing the Power of Adversarial Prompting and Large Language Models for Robust Hypothesis Generation in Astronomy 20 Jun 2023 · 0 repositories · arXiv:2306.11648
-
Learning to Generate Better Than Your LLM 20 Jun 2023 · 1 repository · arXiv:2306.11816
-
Surfer: Progressive Reasoning with World Models for Robotic Manipulation 20 Jun 2023 · 0 repositories · arXiv:2306.11335
-
BayLing: Bridging Cross-lingual Alignment and Instruction Following through Interactive Translation for Large Language Models 19 Jun 2023 · 1 repository · arXiv:2306.10968
-
Temporal Data Meets LLM -- Explainable Financial Time Series Forecasting 19 Jun 2023 · 0 repositories · arXiv:2306.11025
-
News Verifiers Showdown: A Comparative Performance Evaluation of ChatGPT 3.5, ChatGPT 4.0, Bing AI, and Bard in News Fact-Checking 18 Jun 2023 · 0 repositories · arXiv:2306.17176
-
Snowman: A Million-scale Chinese Commonsense Knowledge Graph Distilled from Foundation Model 17 Jun 2023 · 0 repositories · arXiv:2306.10241
-
AD-AutoGPT: An Autonomous GPT for Alzheimer's Disease Infodemiology 16 Jun 2023 · 0 repositories · arXiv:2306.10095
-
Is Self-Repair a Silver Bullet for Code Generation? 16 Jun 2023 · 1 repository · arXiv:2306.09896
-
Evaluating Superhuman Models with Consistency Checks 16 Jun 2023 · 2 repositories · arXiv:2306.09983
-
Explaining Legal Concepts with Augmented Large Language Models (GPT-4) 15 Jun 2023 · 0 repositories · arXiv:2306.09525
-
Exploring the MIT Mathematics and EECS Curriculum Using Large Language Models 15 Jun 2023 · 0 repositories · arXiv:2306.08997
-
Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses 15 Jun 2023 · 0 repositories · arXiv:2306.10073
-
Ensembled Prediction Intervals for Causal Outcomes Under Hidden Confounding 15 Jun 2023 · 0 repositories · arXiv:2306.09520
-
arXiVeri: Automatic table verification with GPT 13 Jun 2023 · 1 repository · arXiv:2306.07968
-
Can ChatGPT Enable ITS? The Case of Mixed Traffic Control via Reinforcement Learning 13 Jun 2023 · 1 repository · arXiv:2306.08094
-
h2oGPT: Democratizing Large Language Models 13 Jun 2023 · 2 repositories · arXiv:2306.08161
-
Human-Like Intuitive Behavior and Reasoning Biases Emerged in Language Models -- and Disappeared in GPT-4 13 Jun 2023 · 0 repositories · arXiv:2306.07622
-
XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models 13 Jun 2023 · 1 repository · arXiv:2306.07971
-
Large language models and (non-)linguistic recursion 12 Jun 2023 · 0 repositories · arXiv:2306.07195
-
Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence 12 Jun 2023 · 0 repositories · arXiv:2306.07075
-
Lost in Translation: Large Language Models in Non-English Content Analysis 12 Jun 2023 · 0 repositories · arXiv:2306.07377
-
Prompt-based Extraction of Social Determinants of Health Using Few-shot Learning 12 Jun 2023 · 0 repositories · arXiv:2306.07170
-
Inductive reasoning in humans and large language models 11 Jun 2023 · 1 repository · arXiv:2306.06548
-
14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon 9 Jun 2023 · 2 repositories · arXiv:2306.06283Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena 9 Jun 2023 · 11 repositories · arXiv:2306.05685Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models 8 Jun 2023 · 1 repository · arXiv:2306.05179
-
Mapping the Challenges of HCI: An Application and Evaluation of ChatGPT and GPT-4 for Mining Insights at Scale 8 Jun 2023 · 0 repositories · arXiv:2306.05036
-
ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases 8 Jun 2023 · 3 repositories · arXiv:2306.05301Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Good Data, Large Data, or No Data? Comparing Three Approaches in Developing Research Aspect Classifiers for Biomedical Papers 7 Jun 2023 · 1 repository · arXiv:2306.04820
-
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources 7 Jun 2023 · 4 repositories · arXiv:2306.04751Syntology official (archive's flag): 5 ran · 6 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models 7 Jun 2023 · 2 repositories · arXiv:2306.04757Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
The Two Word Test: A Semantic Benchmark for Large Language Models 7 Jun 2023 · 1 repository · arXiv:2306.04610
-
Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset 5 Jun 2023 · 1 repository · arXiv:2306.03030Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Orca: Progressive Learning from Complex Explanation Traces of GPT-4 5 Jun 2023 · 4 repositories · arXiv:2306.02707
-
SelfEvolve: A Code Evolution Framework via Large Language Models 5 Jun 2023 · 0 repositories · arXiv:2306.02907
-
Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions 4 Jun 2023 · 1 repository · arXiv:2306.02224Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
MathChat: Converse to Tackle Challenging Math Problems with LLM Agents 2 Jun 2023 · 2 repositories · arXiv:2306.01337
-
Can LLMs like GPT-4 outperform traditional AI tools in dementia diagnosis? Maybe, but not today 2 Jun 2023 · 0 repositories · arXiv:2306.01499
-
Evaluating Language Models for Mathematics through Interactions 2 Jun 2023 · 1 repository · arXiv:2306.01694Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Hybrid Long Document Summarization using C2F-FAR and ChatGPT: A Practical Study 1 Jun 2023 · 0 repositories · arXiv:2306.01169
-
LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day 1 Jun 2023 · 1 repository · arXiv:2306.00890
-
ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing 1 Jun 2023 · 0 repositories · arXiv:2306.00622
-
Automated Annotation with Generative AI Requires Validation 31 May 2023 · 0 repositories · arXiv:2306.00176
-
Evaluating GPT's Programming Capability through CodeWars' Katas 31 May 2023 · 0 repositories · arXiv:2306.01784
-
Scaling Evidence-based Instructional Design Expertise through Large Language Models 31 May 2023 · 0 repositories · arXiv:2306.01006
-
AlphaBlock: Embodied Finetuning for Vision-Language Reasoning in Robot Manipulation 30 May 2023 · 0 repositories · arXiv:2305.18898
-
Does Conceptual Representation Require Embodiment? Insights From Large Language Models 30 May 2023 · 0 repositories · arXiv:2305.19103
-
GPT4GEO: How a Language Model Sees the World's Geography 30 May 2023 · 0 repositories · arXiv:2306.00020
-
GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction 30 May 2023 · 1 repository · arXiv:2305.18752
-
Self-Verification Improves Few-Shot Clinical Information Extraction 30 May 2023 · 1 repository · arXiv:2306.00024Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Chatbots to ChatGPT in a Cybersecurity Space: Evolution, Vulnerabilities, Attacks, Challenges, and Future Recommendations 29 May 2023 · 0 repositories · arXiv:2306.09255
-
Controllable Text-to-Image Generation with GPT-4 29 May 2023 · 0 repositories · arXiv:2305.18583
-
Do Language Models Know When They're Hallucinating References? 29 May 2023 · 1 repository · arXiv:2305.18248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Game of Tones: Faculty detection of GPT-4 generated content in university assessments 29 May 2023 · 0 repositories · arXiv:2305.18081
-
Large Language Models are not Fair Evaluators 29 May 2023 · 1 repository · arXiv:2305.17926Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models 29 May 2023 · 1 repository · arXiv:2305.18189Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)