Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 12
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 12 of 29: papers 1,101 to 1,200 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Grading Massive Open Online Courses Using Large Language Models 16 Jun 2024 · 0 repositories · arXiv:2406.11102
-
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars 16 Jun 2024 · 1 repository · arXiv:2406.11035
-
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences 16 Jun 2024 · 0 repositories · arXiv:2406.11069
-
Beyond Raw Videos: Understanding Edited Videos with Large Multimodal Model 15 Jun 2024 · 1 repository · arXiv:2406.10484
-
Automating Pharmacovigilance Evidence Generation: Using Large Language Models to Produce Context-Aware SQL 15 Jun 2024 · 0 repositories · arXiv:2406.10690
-
SparseCL: Sparse Contrastive Learning for Contradiction Retrieval 15 Jun 2024 · 0 repositories · arXiv:2406.10746
-
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages 14 Jun 2024 · 1 repository · arXiv:2406.09948Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Bootstrapping Language Models with DPO Implicit Rewards 14 Jun 2024 · 1 repository · arXiv:2406.09760Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Domain-Specific Shorthand for Generation Based on Context-Free Grammar 14 Jun 2024 · 0 repositories · arXiv:2406.10442
-
Evaluating LLM-driven User-Intent Formalization for Verification-Aware Languages 14 Jun 2024 · 0 repositories · arXiv:2406.09757
-
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 14 Jun 2024 · 0 repositories · arXiv:2406.09781
-
Know the Unknown: An Uncertainty-Sensitive Method for LLM Instruction Tuning 14 Jun 2024 · 1 repository · arXiv:2406.10099Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Neural Concept Binder 14 Jun 2024 · 1 repository · arXiv:2406.09949Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models 13 Jun 2024 · 1 repository · arXiv:2406.09321Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination 13 Jun 2024 · 0 repositories · arXiv:2406.08818
-
OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning 13 Jun 2024 · 0 repositories · arXiv:2406.08858
-
ReadCtrl: Personalizing text generation with readability-controlled instruction learning 13 Jun 2024 · 0 repositories · arXiv:2406.09205
-
A Sociotechnical Lens for Evaluating Computer Vision Models: A Case Study on Detecting and Reasoning about Gender and Emotion 12 Jun 2024 · 0 repositories · arXiv:2406.08222
-
Automated Information Extraction from Thyroid Operation Narrative: A Comparative Study of GPT-4 and Fine-tuned KoELECTRA 12 Jun 2024 · 0 repositories · arXiv:2406.07922
-
DafnyBench: A Benchmark for Formal Software Verification 12 Jun 2024 · 1 repository · arXiv:2406.08467
-
HelpSteer2: Open-source dataset for training top-performing reward models 12 Jun 2024 · 1 repository · arXiv:2406.08673
-
How well it works: Benchmarking performance of GPT models on medical natural language processing tasks 12 Jun 2024 · 0 repositories
-
Making Task-Oriented Dialogue Datasets More Natural by Synthetically Generating Indirect User Requests 12 Jun 2024 · 0 repositories · arXiv:2406.07794
-
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge 12 Jun 2024 · 1 repository · arXiv:2406.07791Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
Mistral-C2F: Coarse to Fine Actor for Analytical and Reasoning Enhancement in RLHF and Effective-Merged LLMs 12 Jun 2024 · 0 repositories · arXiv:2406.08657
-
Supportiveness-based Knowledge Rewriting for Retrieval-augmented Language Modeling 12 Jun 2024 · 0 repositories · arXiv:2406.08116
-
Tailoring Generative AI Chatbots for Multiethnic Communities in Disaster Preparedness Communication: Extending the CASA Paradigm 12 Jun 2024 · 1 repository · arXiv:2406.08411
-
VeraCT Scan: Retrieval-Augmented Fake News Detection with Justifiable Reasoning 12 Jun 2024 · 0 repositories · arXiv:2406.10289
-
What If We Recaption Billions of Web Images with LLaMA-3? 12 Jun 2024 · 0 repositories · arXiv:2406.08478
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations 11 Jun 2024 · 1 repository · arXiv:2406.07358Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Beyond Words: On Large Language Models Actionability in Mission-Critical Risk Analysis 11 Jun 2024 · 0 repositories · arXiv:2406.10273
-
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs 11 Jun 2024 · 1 repository · arXiv:2406.07080Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models 11 Jun 2024 · 1 repository · arXiv:2406.07594Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Validating LLM-Generated Programs with Metamorphic Prompt Testing 11 Jun 2024 · 0 repositories · arXiv:2406.06864
-
Annotation alignment: Comparing LLM and human annotations of conversational safety 10 Jun 2024 · 0 repositories · arXiv:2406.06369
-
Can Language Models Serve as Text-Based World Simulators? 10 Jun 2024 · 0 repositories · arXiv:2406.06485
-
Data-Efficient Learning with Neural Programs 10 Jun 2024 · 1 repository · arXiv:2406.06246Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning 10 Jun 2024 · 1 repository · arXiv:2406.06469Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
In-Context Learning and Fine-Tuning GPT for Argument Mining 10 Jun 2024 · 1 repository · arXiv:2406.06699
-
SecureNet: A Comparative Study of DeBERTa and Large Language Models for Phishing Detection 10 Jun 2024 · 0 repositories · arXiv:2406.06663
-
A Knowledge-Component-Based Methodology for Evaluating AI Assistants 9 Jun 2024 · 0 repositories · arXiv:2406.05603
-
Are Large Language Models Actually Good at Text Style Transfer? 9 Jun 2024 · 1 repository · arXiv:2406.05885
-
Exploring the Efficacy of Large Language Models (GPT-4) in Binary Reverse Engineering 9 Jun 2024 · 0 repositories · arXiv:2406.06637
-
Large Language Models Memorize Sensor Datasets! Implications on Human Activity Recognition Research 9 Jun 2024 · 0 repositories · arXiv:2406.05900
-
Text2VP: Generative AI for Visual Programming and Parametric Modeling 9 Jun 2024 · 0 repositories · arXiv:2407.07732
-
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding 8 Jun 2024 · 1 repository · arXiv:2406.05540Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Do LLMs Recognize me, When I is not me: Assessment of LLMs Understanding of Turkish Indexical Pronouns in Indexical Shift Contexts 8 Jun 2024 · 0 repositories · arXiv:2406.05569
-
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner 8 Jun 2024 · 0 repositories · arXiv:2406.05498
-
Teaching-Assistant-in-the-Loop: Improving Knowledge Distillation from Imperfect Teacher Models in Low-Budget Scenarios 8 Jun 2024 · 0 repositories · arXiv:2406.05322
-
Toward Reliable Ad-hoc Scientific Information Extraction: A Case Study on Two Materials Datasets 8 Jun 2024 · 1 repository · arXiv:2406.05348
-
Are Large Language Models More Empathetic than Humans? 7 Jun 2024 · 0 repositories · arXiv:2406.05063
-
Diving Deep into the Motion Representation of Video-Text Models 7 Jun 2024 · 1 repository · arXiv:2406.05075
-
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents 7 Jun 2024 · 2 repositories · arXiv:2406.06613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Hints-In-Browser: Benchmarking Language Models for Programming Feedback Generation 7 Jun 2024 · 0 repositories · arXiv:2406.05053
-
DALD: Improving Logits-based Detector without Logits from Black-box LLMs 7 Jun 2024 · 1 repository · arXiv:2406.05232
-
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models 7 Jun 2024 · 1 repository · arXiv:2406.05113Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs 7 Jun 2024 · 1 repository · arXiv:2406.05194
-
Low-Resource Cross-Lingual Summarization through Few-Shot Learning with Large Language Models 7 Jun 2024 · 0 repositories · arXiv:2406.04630
-
Mixture-of-Agents Enhances Large Language Model Capabilities 7 Jun 2024 · 3 repositories · arXiv:2406.04692Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
Benchmark Data Contamination of Large Language Models: A Survey 6 Jun 2024 · 0 repositories · arXiv:2406.04244
-
Characterizing Similarities and Divergences in Conversational Tones in Humans and LLMs by Sampling with People 6 Jun 2024 · 1 repository · arXiv:2406.04278
-
Exploring the Latest LLMs for Leaderboard Extraction 6 Jun 2024 · 0 repositories · arXiv:2406.04383
-
Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning 6 Jun 2024 · 1 repository · arXiv:2406.03718
-
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning 6 Jun 2024 · 0 repositories · arXiv:2406.04520
-
RoboCoder: Robotic Learning from Basic Skills to General Tasks with Large Language Models 6 Jun 2024 · 0 repositories · arXiv:2406.03757
-
Scaling and evaluating sparse autoencoders 6 Jun 2024 · 5 repositories · arXiv:2406.04093Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
Tool-Planner: Task Planning with Clusters across Multiple Tools 6 Jun 2024 · 1 repository · arXiv:2406.03807Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
UltraMedical: Building Specialized Generalists in Biomedicine 6 Jun 2024 · 1 repository · arXiv:2406.03949Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Analyzing LLM Behavior in Dialogue Summarization: Unveiling Circumstantial Hallucination Trends 5 Jun 2024 · 0 repositories · arXiv:2406.03487
-
BIPED: Pedagogically Informed Tutoring System for ESL Education 5 Jun 2024 · 0 repositories · arXiv:2406.03486
-
Evaluating the Efficacy of Large Language Models in Detecting Fake News: A Comparative Analysis 5 Jun 2024 · 0 repositories · arXiv:2406.06584
-
From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation 5 Jun 2024 · 0 repositories · arXiv:2406.03030
-
The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games 5 Jun 2024 · 0 repositories · arXiv:2406.03299
-
What is the Best Way for ChatGPT to Translate Poetry? 5 Jun 2024 · 0 repositories · arXiv:2406.03450
-
Dishonesty in Helpful and Harmless Alignment 4 Jun 2024 · 0 repositories · arXiv:2406.01931
-
Eliciting the Priors of Large Language Models using Iterated In-Context Learning 4 Jun 2024 · 0 repositories · arXiv:2406.01860
-
Large Language Model-Enabled Multi-Agent Manufacturing Systems 4 Jun 2024 · 0 repositories · arXiv:2406.01893
-
Multiple Choice Questions and Large Languages Models: A Case Study with Fictional Medical Data 4 Jun 2024 · 1 repository · arXiv:2406.02394
-
BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models 3 Jun 2024 · 0 repositories · arXiv:2406.00083
-
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models 3 Jun 2024 · 1 repository · arXiv:2406.01698
-
Predicting Drug-Gene Relations via Analogy Tasks with Word Embeddings 3 Jun 2024 · 1 repository · arXiv:2406.00984
-
Superhuman performance in urology board questions by an explainable large language model enabled for context integration of the European Association of Urology guidelines: the UroBot study 3 Jun 2024 · 0 repositories · arXiv:2406.01428
-
Utilizing Large Language Models for Automating Technical Customer Support 3 Jun 2024 · 0 repositories · arXiv:2406.01407
-
An Early Investigation into the Utility of Multimodal Large Language Models in Medical Imaging 2 Jun 2024 · 0 repositories · arXiv:2406.00667
-
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction 2 Jun 2024 · 1 repository · arXiv:2406.00755Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Presence or Absence: Are Unknown Word Usages in Dictionaries? 2 Jun 2024 · 1 repository · arXiv:2406.00656
-
An Evaluation Benchmark for Autoformalization in Lean4 1 Jun 2024 · 0 repositories · arXiv:2406.06555
-
Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios 1 Jun 2024 · 0 repositories · arXiv:2406.00343
-
Phased Instruction Fine-Tuning for Large Language Models 1 Jun 2024 · 1 repository · arXiv:2406.04371Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Leveraging Large Language Models for Entity Matching 31 May 2024 · 0 repositories · arXiv:2405.20624
-
Query2CAD: Generating CAD models using natural language queries 31 May 2024 · 1 repository · arXiv:2406.00144Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Superlatives in Context: Modeling the Implicit Semantics of Superlatives 31 May 2024 · 1 repository · arXiv:2405.20967
-
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 31 May 2024 · 1 repository · arXiv:2405.21075Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
An Automatic Question Usability Evaluation Toolkit 30 May 2024 · 1 repository · arXiv:2405.20529
-
ANAH: Analytical Annotation of Hallucinations in Large Language Models 30 May 2024 · 1 repository · arXiv:2405.20315Syntology official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
AutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization 30 May 2024 · 0 repositories · arXiv:2405.19668
-
Automated Generation and Tagging of Knowledge Components from Multiple-Choice Questions 30 May 2024 · 1 repository · arXiv:2405.20526
-
Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation 30 May 2024 · 0 repositories · arXiv:2405.20092
-
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning 30 May 2024 · 1 repository · arXiv:2405.20139Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
LLaMEA: A Large Language Model Evolutionary Algorithm for Automatically Generating Metaheuristics 30 May 2024 · 2 repositories · arXiv:2405.20132Syntology official (archive's flag): 3 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)