Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 13
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 13 of 29: papers 1,201 to 1,300 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
PATIENT-Ψ: Using Large Language Models to Simulate Patients for Training Mental Health Professionals 30 May 2024 · 1 repository · arXiv:2405.19660Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations 30 May 2024 · 1 repository · arXiv:2405.19740Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Phantom: General Trigger Attacks on Retrieval Augmented Language Generation 30 May 2024 · 0 repositories · arXiv:2405.20485
-
Preference Alignment with Flow Matching 30 May 2024 · 1 repository · arXiv:2405.19806Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Are You Sure? Rank Them Again: Repeated Ranking For Better Preference Datasets 29 May 2024 · 0 repositories · arXiv:2405.18952
-
Beyond Agreement: Diagnosing the Rationale Alignment of Automated Essay Scoring Methods based on Linguistically-informed Counterfactuals 29 May 2024 · 1 repository · arXiv:2405.19433
-
LLM-based Hierarchical Concept Decomposition for Interpretable Fine-Grained Image Classification 29 May 2024 · 0 repositories · arXiv:2405.18672
-
LLMs achieve adult human performance on higher-order theory of mind tasks 29 May 2024 · 0 repositories · arXiv:2405.18870
-
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications 29 May 2024 · 1 repository · arXiv:2405.19266Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Reverse Image Retrieval Cues Parametric Memory in Multimodal LLMs 29 May 2024 · 1 repository · arXiv:2405.18740
-
Two-Layer Retrieval-Augmented Generation Framework for Low-Resource Medical Question Answering Using Reddit Data: Proof-of-Concept Study 29 May 2024 · 0 repositories · arXiv:2405.19519
-
Aligning to Thousands of Preferences via System Message Generalization 28 May 2024 · 1 repository · arXiv:2405.17977Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
An Empirical Analysis on Large Language Models in Debate Evaluation 28 May 2024 · 1 repository · arXiv:2406.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Benchmarks Underestimate the Readiness of Multi-lingual Dialogue Agents 28 May 2024 · 0 repositories · arXiv:2405.17840
-
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints 28 May 2024 · 0 repositories · arXiv:2405.18028
-
Multi-objective Representation for Numbers in Clinical Narratives: A CamemBERT-Bio-Based Alternative to Large-Scale LLMs 28 May 2024 · 0 repositories · arXiv:2405.18448
-
Notes on Applicability of GPT-4 to Document Understanding 28 May 2024 · 0 repositories · arXiv:2405.18433
-
ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling 28 May 2024 · 1 repository · arXiv:2405.17743Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Peering into the Mind of Language Models: An Approach for Attribution in Contextual Question Answering 28 May 2024 · 1 repository · arXiv:2405.17980
-
Proof of Quality: A Costless Paradigm for Trustless Generative AI Model Inference on Blockchains 28 May 2024 · 0 repositories · arXiv:2405.17934
-
RealitySummary: Exploring On-Demand Mixed Reality Text Summarization and Question Answering using Large Language Models 28 May 2024 · 0 repositories · arXiv:2405.18620
-
Thai Winograd Schemas: A Benchmark for Thai Commonsense Reasoning 28 May 2024 · 1 repository · arXiv:2405.18375
-
The Battle of LLMs: A Comparative Study in Conversational QA Tasks 28 May 2024 · 0 repositories · arXiv:2405.18344
-
Autoformalizing Euclidean Geometry 27 May 2024 · 1 repository · arXiv:2405.17216Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 2 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
CHESS: Contextual Harnessing for Efficient SQL Synthesis 27 May 2024 · 2 repositories · arXiv:2405.16755Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Cost-efficient Knowledge-based Question Answering with Large Language Models 27 May 2024 · 0 repositories · arXiv:2405.17337
-
Interesting Scientific Idea Generation using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders 27 May 2024 · 1 repository · arXiv:2405.17044
-
REVECA: Adaptive Planning and Trajectory-based Validation in Cooperative Language Agents using Information Relevance and Relative Proximity 27 May 2024 · 0 repositories · arXiv:2405.16751
-
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs 27 May 2024 · 1 repository · arXiv:2405.17013Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
ReflectionCoder: Learning from Reflection Sequence for Enhanced One-off Code Generation 27 May 2024 · 1 repository · arXiv:2405.17057Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
RTL-Repo: A Benchmark for Evaluating LLMs on Large-Scale RTL Design Projects 27 May 2024 · 1 repository · arXiv:2405.17378
-
Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models 27 May 2024 · 1 repository · arXiv:2405.16833Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
THREAD: Thinking Deeper with Recursive Spawning 27 May 2024 · 1 repository · arXiv:2405.17402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
DarijaBanking: A New Resource for Overcoming Language Barriers in Banking Intent Detection for Moroccan Arabic Speakers 26 May 2024 · 1 repository · arXiv:2405.16482
-
Planning with Multi-Constraints via Collaborative Language Agents 26 May 2024 · 1 repository · arXiv:2405.16510
-
Comparative Analysis of Open-Source Language Models in Summarizing Medical Text Data 25 May 2024 · 0 repositories · arXiv:2405.16295
-
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models 25 May 2024 · 1 repository · arXiv:2405.16282Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; the one sample that ran constructed an object rather than computing a result (of 6 harvested samples)
-
GeneAgent: Self-verification Language Agent for Gene Set Knowledge Discovery using Domain Databases 25 May 2024 · 0 repositories · arXiv:2405.16205
-
HETHUB: A Distributed Training System with Heterogeneous Cluster for Large-Scale Models 25 May 2024 · 0 repositories · arXiv:2405.16256
-
Picturing Ambiguity: A Visual Twist on the Winograd Schema Challenge 25 May 2024 · 0 repositories · arXiv:2405.16277
-
STRIDE: A Tool-Assisted LLM Agent Framework for Strategic and Interactive Decision-Making 25 May 2024 · 1 repository · arXiv:2405.16376Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
An Evaluation of Estimative Uncertainty in Large Language Models 24 May 2024 · 0 repositories · arXiv:2405.15185
-
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation 24 May 2024 · 1 repository · arXiv:2405.15307Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving 24 May 2024 · 1 repository · arXiv:2405.15324
-
CulturePark: Boosting Cross-cultural Understanding in Large Language Models 24 May 2024 · 1 repository · arXiv:2405.15145Syntology official: harvested, nothing ran · 0 ran · 7 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Large Language Models Reflect Human Citation Patterns with a Heightened Citation Bias 24 May 2024 · 1 repository · arXiv:2405.15739
-
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models 24 May 2024 · 0 repositories · arXiv:2405.15936
-
A Declarative System for Optimizing AI Workloads 23 May 2024 · 1 repository · arXiv:2405.14696
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
AutoCoder: Enhancing Code Large Language Model with AIEV-Instruct 23 May 2024 · 1 repository · arXiv:2405.14906
-
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 23 May 2024 · 0 repositories · arXiv:2405.14333
-
Efficient Medical Question Answering with Knowledge-Augmented Question Generation 23 May 2024 · 1 repository · arXiv:2405.14654
-
Evaluating Large Language Models for Public Health Classification and Extraction Tasks 23 May 2024 · 0 repositories · arXiv:2405.14766
-
Exploring the use of a Large Language Model for data extraction in systematic reviews: a rapid feasibility study 23 May 2024 · 0 repositories · arXiv:2405.14445
-
Impact of Non-Standard Unicode Characters on Security and Comprehension in Large Language Models 23 May 2024 · 1 repository · arXiv:2405.14490
-
Improving Language Models Trained on Translated Data with Continual Pre-Training and Dictionary Learning Analysis 23 May 2024 · 0 repositories · arXiv:2405.14277
-
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models 23 May 2024 · 1 repository · arXiv:2405.14365Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Perception of Knowledge Boundary for Large Language Models through Semi-open-ended Question Answering 23 May 2024 · 0 repositories · arXiv:2405.14383
-
Evaluating Large Language Models with Human Feedback: Establishing a Swedish Benchmark 22 May 2024 · 1 repository · arXiv:2405.14006
-
Why Not Transform Chat Large Language Models to Non-English? 22 May 2024 · 1 repository · arXiv:2405.13923
-
WordGame: Efficient & Effective LLM Jailbreak via Simultaneous Obfuscation in Query and Response 22 May 2024 · 0 repositories · arXiv:2405.14023
-
BiomedParse: a biomedical foundation model for image parsing of everything everywhere all at once 21 May 2024 · 0 repositories · arXiv:2405.12971
-
Generative AI in Cybersecurity: A Comprehensive Review of LLM Applications and Vulnerabilities 21 May 2024 · 0 repositories · arXiv:2405.12750
-
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation 21 May 2024 · 0 repositories · arXiv:2405.13077
-
PathOCL: Path-Based Prompt Augmentation for OCL Generation with GPT-4 21 May 2024 · 0 repositories · arXiv:2405.12450
-
Can AI Relate: Testing Large Language Model Response for Mental Health Support 20 May 2024 · 1 repository · arXiv:2405.12021Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
CT-Eval: Benchmarking Chinese Text-to-Table Performance in Large Language Models 20 May 2024 · 0 repositories · arXiv:2405.12174
-
Fennec: Fine-grained Language Model Evaluation and Correction Extended through Branching and Bridging 20 May 2024 · 1 repository · arXiv:2405.12163
-
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving 20 May 2024 · 0 repositories · arXiv:2405.12205
-
Hummer: Towards Limited Competitive Preference Dataset 19 May 2024 · 0 repositories · arXiv:2405.11647
-
Large Language Models Can Infer Personality from Free-Form User Interactions 19 May 2024 · 0 repositories · arXiv:2405.13052
-
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation 19 May 2024 · 1 repository · arXiv:2405.11430
-
Automating PTSD Diagnostics in Clinical Interviews: Leveraging Large Language Models for Trauma Assessments 18 May 2024 · 0 repositories · arXiv:2405.11178
-
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ? 18 May 2024 · 0 repositories · arXiv:2405.11407
-
ActiveLLM: Large Language Model-based Active Learning for Textual Few-Shot Scenarios 17 May 2024 · 0 repositories · arXiv:2405.10808
-
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations 17 May 2024 · 0 repositories · arXiv:2405.11100
-
Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset 17 May 2024 · 2 repositories · arXiv:2405.10542
-
Enhancing Dialogue State Tracking Models through LLM-backed User-Agents Simulation 17 May 2024 · 0 repositories · arXiv:2405.13037
-
Evaluation of large language model performance on the Biomedical Language Understanding and Reasoning Benchmark 17 May 2024 · 0 repositories
-
Language Models can Evaluate Themselves via Probability Discrepancy 17 May 2024 · 1 repository · arXiv:2405.10516
-
Large Language Models in Wireless Application Design: In-Context Learning-enhanced Automatic Network Intrusion Detection 17 May 2024 · 0 repositories · arXiv:2405.11002
-
Observational Scaling Laws and the Predictability of Language Model Performance 17 May 2024 · 1 repository · arXiv:2405.10938Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Dynamic In-context Learning with Conversational Models for Data Extraction and Materials Property Prediction 16 May 2024 · 1 repository · arXiv:2405.10448
-
FinTextQA: A Dataset for Long-form Financial Question Answering 16 May 2024 · 0 repositories · arXiv:2405.09980
-
The AI Collaborator: Bridging Human-AI Interaction in Educational and Professional Settings 16 May 2024 · 0 repositories · arXiv:2405.10460
-
Transcript of GPT-4 playing a rogue AGI in a Matrix Game 16 May 2024 · 0 repositories · arXiv:2405.10997
-
Comparing the Efficacy of GPT-4 and Chat-GPT in Mental Health Care: A Blind Assessment of Large Language Models for Psychological Support 15 May 2024 · 0 repositories · arXiv:2405.09300
-
Exploring the Potential of Large Language Models for Automation in Technical Customer Service 15 May 2024 · 0 repositories · arXiv:2405.09161
-
Intelligent Tutor: Leveraging ChatGPT and Microsoft Copilot Studio to Deliver a Generative AI Student Support and Feedback System within Teams 15 May 2024 · 0 repositories · arXiv:2405.13024
-
Simulating Policy Impacts: Developing a Generative Scenario Writing Method to Evaluate the Perceived Effects of Regulation 15 May 2024 · 0 repositories · arXiv:2405.09679
-
SQL-to-Schema Enhances Schema Linking in Text-to-SQL 15 May 2024 · 0 repositories · arXiv:2405.09593
-
Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models 15 May 2024 · 1 repository · arXiv:2405.09454
-
Word Alignment as Preference for Machine Translation 15 May 2024 · 0 repositories · arXiv:2405.09223
-
A Comprehensive Survey of Large Language Models and Multimodal Large Language Models in Medicine 14 May 2024 · 0 repositories · arXiv:2405.08603
-
Can Language Models Explain Their Own Classification Behavior? 13 May 2024 · 1 repository · arXiv:2405.07436
-
Coding historical causes of death data with Large Language Models 13 May 2024 · 1 repository · arXiv:2405.07560
-
MetaReflection: Learning Instructions for Language Agents using Past Reflections 13 May 2024 · 0 repositories · arXiv:2405.13009
-
Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots 13 May 2024 · 0 repositories · arXiv:2405.07990
-
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing 13 May 2024 · 1 repository · arXiv:2405.07726Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
SambaNova SN40L: Scaling the AI Memory Wall with Dataflow and Composition of Experts 13 May 2024 · 0 repositories · arXiv:2405.07518