Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 11
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 11 of 29: papers 1,001 to 1,100 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Pron vs Prompt: Can Large Language Models already Challenge a World-Class Fiction Author at Creative Text Writing? 1 Jul 2024 · 0 repositories · arXiv:2407.01119
-
Roleplay-doh: Enabling Domain-Experts to Create LLM-simulated Patients via Eliciting and Adhering to Principles 1 Jul 2024 · 0 repositories · arXiv:2407.00870
-
Evaluation of Bias Towards Medical Professionals in Large Language Models 30 Jun 2024 · 0 repositories · arXiv:2407.12031
-
LLM-Generated Natural Language Meets Scaling Laws: New Explorations and Data Augmentation Methods 29 Jun 2024 · 0 repositories · arXiv:2407.00322
-
Too Late to Train, Too Early To Use? A Study on Necessity and Viability of Low-Resource Bengali LLMs 29 Jun 2024 · 0 repositories · arXiv:2407.00416
-
Urban Visual Appeal According to ChatGPT: Contrasting AI and Human Insights 29 Jun 2024 · 0 repositories · arXiv:2407.14268
-
AnomaLLMy -- Detecting anomalous tokens in black-box LLMs through low-confidence single-token predictions 28 Jun 2024 · 1 repository · arXiv:2406.19840
-
Can GPT-4 Help Detect Quit Vaping Intentions? An Exploration of Automatic Data Annotation Approach 28 Jun 2024 · 0 repositories · arXiv:2407.00167
-
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation 28 Jun 2024 · 0 repositories · arXiv:2406.20053
-
Can Large Language Models Generate High-quality Patent Claims? 27 Jun 2024 · 1 repository · arXiv:2406.19465
-
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets 27 Jun 2024 · 1 repository · arXiv:2406.18906Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
The Model Arena for Cross-lingual Sentiment Analysis: A Comparative Study in the Era of Large Language Models 27 Jun 2024 · 0 repositories · arXiv:2406.19358
-
UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models 27 Jun 2024 · 1 repository · arXiv:2406.18966Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Adversarial Search Engine Optimization for Large Language Models 26 Jun 2024 · 0 repositories · arXiv:2406.18382
-
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets 26 Jun 2024 · 0 repositories · arXiv:2406.18518
-
BADGE: BADminton report Generation and Evaluation with LLM 26 Jun 2024 · 1 repository · arXiv:2406.18116
-
Improving Entity Recognition Using Ensembles of Deep Learning and Fine-tuned Large Language Models: A Case Study on Adverse Event Extraction from Multiple Sources 26 Jun 2024 · 0 repositories · arXiv:2406.18049
-
Jailbreaking LLMs with Arabic Transliteration and Arabizi 26 Jun 2024 · 1 repository · arXiv:2406.18725Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Octo-planner: On-device Language Model for Planner-Action Agents 26 Jun 2024 · 0 repositories · arXiv:2406.18082
-
Re-Ranking Step by Step: Investigating Pre-Filtering for Re-Ranking with Large Language Models 26 Jun 2024 · 0 repositories · arXiv:2406.18740
-
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 26 Jun 2024 · 1 repository · arXiv:2406.18629Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability 26 Jun 2024 · 1 repository · arXiv:2406.18365Syntology official (archive's flag): 5 ran · 5 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs 26 Jun 2024 · 4 repositories · arXiv:2406.18495Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Accelerating Clinical Evidence Synthesis with Large Language Models 25 Jun 2024 · 0 repositories · arXiv:2406.17755
-
ARES: Alternating Reinforcement Learning and Supervised Fine-Tuning for Enhanced Multi-Modal Chain-of-Thought Reasoning Through Diverse AI Feedback 25 Jun 2024 · 1 repository · arXiv:2407.00087
-
Autonomous Prompt Engineering in Large Language Models 25 Jun 2024 · 0 repositories · arXiv:2407.11000
-
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels 25 Jun 2024 · 0 repositories · arXiv:2406.17633
-
LongIns: A Challenging Long-context Instruction-based Exam for LLMs 25 Jun 2024 · 0 repositories · arXiv:2406.17588
-
Anomaly Detection of Tabular Data Using LLMs 24 Jun 2024 · 0 repositories · arXiv:2406.16308
-
Classification of Geological Borehole Descriptions Using a Domain Adapted Large Language Model 24 Jun 2024 · 0 repositories · arXiv:2407.10991
-
Exploring Factual Entailment with NLI: A News Media Study 24 Jun 2024 · 0 repositories · arXiv:2406.16842
-
Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders 24 Jun 2024 · 0 repositories · arXiv:2406.16510
-
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models 24 Jun 2024 · 1 repository · arXiv:2406.17169
-
PlagBench: Exploring the Duality of Large Language Models in Plagiarism Generation and Detection 24 Jun 2024 · 0 repositories · arXiv:2406.16288
-
UNO Arena for Evaluating Sequential Decision-Making Capability of Large Language Models 24 Jun 2024 · 0 repositories · arXiv:2406.16382
-
USDC: A Dataset of User Stance and Dogmatism in Long Conversations 24 Jun 2024 · 0 repositories · arXiv:2406.16833
-
Enhancing Commentary Strategies for Imperfect Information Card Games: A Study of Large Language Models in Guandan Commentary 23 Jun 2024 · 1 repository · arXiv:2406.17807
-
GraphEval2000: Benchmarking and Improving Large Language Models on Graph Datasets 23 Jun 2024 · 0 repositories · arXiv:2406.16176
-
Can LLMs Generate Visualizations with Dataless Prompts? 22 Jun 2024 · 0 repositories · arXiv:2406.17805
-
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level 22 Jun 2024 · 3 repositories · arXiv:2406.15741Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
A GPT-based Code Review System for Programming Language Learning 21 Jun 2024 · 0 repositories · arXiv:2407.04722
-
A SMART Mnemonic Sounds like "Glue Tonic": Mixing LLMs with Student Feedback to Make Mnemonic Learning Stick 21 Jun 2024 · 1 repository · arXiv:2406.15352
-
Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling 21 Jun 2024 · 0 repositories · arXiv:2406.15527
-
Efficient Continual Pre-training by Mitigating the Stability Gap 21 Jun 2024 · 0 repositories · arXiv:2406.14833
-
ESC-Eval: Evaluating Emotion Support Conversations in Large Language Models 21 Jun 2024 · 3 repositories · arXiv:2406.14952Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
How Effective is GPT-4 Turbo in Generating School-Level Questions from Textbooks Based on Bloom's Revised Taxonomy? 21 Jun 2024 · 0 repositories · arXiv:2406.15211
-
InternLM-Law: An Open Source Chinese Legal Large Language Model 21 Jun 2024 · 1 repository · arXiv:2406.14887
-
TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings 21 Jun 2024 · 1 repository · arXiv:2406.15586
-
V-RECS, a Low-Cost LLM4VIS Recommender with Explanations, Captioning and Suggestions 21 Jun 2024 · 1 repository · arXiv:2406.15259
-
What Teaches Robots to Walk, Teaches Them to Trade too -- Regime Adaptive Execution using Informed Data and LLMs 20 Jun 2024 · 0 repositories · arXiv:2406.15508
-
A Large Language Model Outperforms Other Computational Approaches to the High-Throughput Phenotyping of Physician Notes 20 Jun 2024 · 0 repositories · arXiv:2406.14757
-
ChatGPT as Research Scientist: Probing GPT's Capabilities as a Research Librarian, Research Ethicist, Data Generator and Data Predictor 20 Jun 2024 · 0 repositories · arXiv:2406.14765
-
CryptoGPT: a 7B model rivaling GPT-4 in the task of analyzing and classifying real-time financial news 20 Jun 2024 · 0 repositories · arXiv:2406.14039
-
DIRAS: Efficient LLM Annotation of Document Relevance in Retrieval Augmented Generation 20 Jun 2024 · 1 repository · arXiv:2406.14162
-
Enhancing the LLM-Based Robot Manipulation Through Human-Robot Collaboration 20 Jun 2024 · 0 repositories · arXiv:2406.14097
-
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions 20 Jun 2024 · 0 repositories · arXiv:2406.13903
-
Identifying User Goals from UI Trajectories 20 Jun 2024 · 0 repositories · arXiv:2406.14314
-
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 20 Jun 2024 · 1 repository · arXiv:2406.14515Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
MR-Ben: A Meta-Reasoning Benchmark for Evaluating System-2 Thinking in LLMs 20 Jun 2024 · 0 repositories · arXiv:2406.13975
-
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors 20 Jun 2024 · 1 repository · arXiv:2406.14598
-
SPL: A Socratic Playground for Learning Powered by Large Language Model 20 Jun 2024 · 0 repositories · arXiv:2406.13919
-
The Use of Multimodal Large Language Models to Detect Objects from Thermal Images: Transportation Applications 20 Jun 2024 · 0 repositories · arXiv:2406.13898
-
TTQA-RS- A break-down prompting approach for Multi-hop Table-Text Question Answering with Reasoning and Summarization 20 Jun 2024 · 0 repositories · arXiv:2406.14732
-
Unmasking Database Vulnerabilities: Zero-Knowledge Schema Inference Attacks in Text-to-SQL Systems 20 Jun 2024 · 0 repositories · arXiv:2406.14545
-
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding 19 Jun 2024 · 1 repository · arXiv:2406.13807Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks 19 Jun 2024 · 1 repository · arXiv:2406.13264Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Is GPT-4 conscious? 19 Jun 2024 · 0 repositories · arXiv:2407.09517
-
MoreHopQA: More Than Multi-hop Reasoning 19 Jun 2024 · 1 repository · arXiv:2406.13397
-
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs 18 Jun 2024 · 2 repositories · arXiv:2406.12288Syntology official (archive's flag): 11 ran · 34 ran (of which 0 constructed an object rather than computing a result; 32 with no instrument failure: 0 honoured, 3 violated, 29 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 36 harvested samples) · 12 pointer-only (licence)
-
Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An Empirical Study 18 Jun 2024 · 0 repositories · arXiv:2406.13049
-
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones? 18 Jun 2024 · 1 repository · arXiv:2406.12809Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 4 pointer-only (licence)
-
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools 18 Jun 2024 · 7 repositories · arXiv:2406.12793Syntology official (archive's flag): 4 ran · 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 29 harvested samples) · 1 pointer-only (licence)
-
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 18 Jun 2024 · 1 repository · arXiv:2407.13690Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
Generative Artificial Intelligence-Guided User Studies: An Application for Air Taxi Services 18 Jun 2024 · 0 repositories · arXiv:2406.12296
-
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts 18 Jun 2024 · 2 repositories · arXiv:2406.12845Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges 18 Jun 2024 · 1 repository · arXiv:2406.12624Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Measuring Psychological Depth in Language Models 18 Jun 2024 · 1 repository · arXiv:2406.12680
-
UBENCH: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions 18 Jun 2024 · 1 repository · arXiv:2406.12784
-
Vernacular? I Barely Know Her: Challenges with Style Control and Stereotyping 18 Jun 2024 · 0 repositories · arXiv:2406.12679
-
A Two-dimensional Zero-shot Dialogue State Tracking Evaluation Method using GPT-4 17 Jun 2024 · 1 repository · arXiv:2406.11651
-
Are Large Language Models True Healthcare Jacks-of-All-Trades? Benchmarking Across Health Professions Beyond Physician Exams 17 Jun 2024 · 1 repository · arXiv:2406.11328
-
Beyond Boundaries: Learning a Universal Entity Taxonomy across Datasets and Languages for Open Named Entity Recognition 17 Jun 2024 · 1 repository · arXiv:2406.11192
-
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting 17 Jun 2024 · 0 repositories · arXiv:2406.11661
-
Decoding the Narratives: Analyzing Personal Drug Experiences Shared on Reddit 17 Jun 2024 · 0 repositories · arXiv:2406.12117
-
Enabling robots to follow abstract instructions and complete complex dynamic tasks 17 Jun 2024 · 0 repositories · arXiv:2406.11231
-
FinTruthQA: A Benchmark Dataset for Evaluating the Quality of Financial Information Disclosure 17 Jun 2024 · 1 repository · arXiv:2406.12009Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation 17 Jun 2024 · 1 repository · arXiv:2406.11503Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Problematic Tokens: Tokenizer Bias in Large Language Models 17 Jun 2024 · 1 repository · arXiv:2406.11214
-
Iterative Length-Regularized Direct Preference Optimization: A Case Study on Improving 7B Language Models to GPT-4 Level 17 Jun 2024 · 0 repositories · arXiv:2406.11817
-
Large Language Models and Knowledge Graphs for Astronomical Entity Disambiguation 17 Jun 2024 · 0 repositories · arXiv:2406.11400
-
MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic 17 Jun 2024 · 0 repositories · arXiv:2406.11385
-
Satyrn: A Platform for Analytics Augmented Generation 17 Jun 2024 · 1 repository · arXiv:2406.12069
-
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models 17 Jun 2024 · 0 repositories · arXiv:2406.12146
-
Small Agent Can Also Rock! Empowering Small Language Models as Hallucination Detector 17 Jun 2024 · 1 repository · arXiv:2406.11277
-
Unveiling and Mitigating Bias in Mental Health Analysis with Large Language Models 17 Jun 2024 · 1 repository · arXiv:2406.12033
-
Can LLMs Understand the Implication of Emphasized Sentences in Dialogue? 16 Jun 2024 · 1 repository · arXiv:2406.11065
-
Distilling Opinions at Scale: Incremental Opinion Summarization using XL-OPSUMM 16 Jun 2024 · 0 repositories · arXiv:2406.10886
-
Enhancing Supermarket Robot Interaction: A Multi-Level LLM Conversational Interface for Handling Diverse Customer Intents 16 Jun 2024 · 0 repositories · arXiv:2406.11047
-
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning 16 Jun 2024 · 0 repositories · arXiv:2406.10834
-
Generating Tables from the Parametric Knowledge of Language Models 16 Jun 2024 · 1 repository · arXiv:2406.10922