Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 3
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 3 of 29: papers 201 to 300 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
RESPONSE: Benchmarking the Ability of Language Models to Undertake Commonsense Reasoning in Crisis Situation 14 Mar 2025 · 0 repositories · arXiv:2503.11348
-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 13 Mar 2025 · 1 repository · arXiv:2503.10635Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Advanced Tool Learning and Selection System (ATLASS): A Closed-Loop Framework Using LLM 13 Mar 2025 · 0 repositories · arXiv:2503.10071
-
ChatGPT Encounters Morphing Attack Detection: Zero-Shot MAD with Multi-Modal Large Language Models and General Vision Models 13 Mar 2025 · 0 repositories · arXiv:2503.10937
-
Do I look like a `cat.n.01` to you? A Taxonomy Image Generation Benchmark 13 Mar 2025 · 0 repositories · arXiv:2503.10357
-
Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search 13 Mar 2025 · 0 repositories · arXiv:2503.10619
-
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents 12 Mar 2025 · 1 repository · arXiv:2503.09780Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
An Evaluation of LLMs for Detecting Harmful Computing Terms 12 Mar 2025 · 0 repositories · arXiv:2503.09341
-
How to Protect Yourself from 5G Radiation? Investigating LLM Responses to Implicit Misinformation 12 Mar 2025 · 1 repository · arXiv:2503.09598
-
Un-Straightening Generative AI: How Queer Artists Surface and Challenge the Normativity of Generative AI Models 12 Mar 2025 · 0 repositories · arXiv:2503.09805
-
EFPC: Towards Efficient and Flexible Prompt Compression 11 Mar 2025 · 0 repositories · arXiv:2503.07956
-
External Knowledge Injection for CLIP-Based Class-Incremental Learning 11 Mar 2025 · 3 repositories · arXiv:2503.08510
-
Seeing What's Not There: Spurious Correlation in Multimodal LLMs 11 Mar 2025 · 0 repositories · arXiv:2503.08884
-
Bot Wars Evolved: Orchestrating Competing LLMs in a Counterstrike Against Phone Scams 10 Mar 2025 · 0 repositories · arXiv:2503.07036
-
Exploring Multimodal Perception in Large Language Models Through Perceptual Strength Ratings 10 Mar 2025 · 0 repositories · arXiv:2503.06980
-
Large model enhanced computational ghost imaging 10 Mar 2025 · 1 repository · arXiv:2503.08710
-
SKG-LLM: Developing a Mathematical Model for Stroke Knowledge Graph Construction Using Large Language Models 9 Mar 2025 · 0 repositories · arXiv:2503.06475
-
Optimizing Generative AI's Accuracy and Transparency in Inductive Thematic Analysis: A Human-AI Comparison 8 Mar 2025 · 0 repositories · arXiv:2503.16485
-
DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL 6 Mar 2025 · 0 repositories · arXiv:2503.04959
-
Incentivizing Multi-Tenant Split Federated Learning for Foundation Models at the Network Edge 6 Mar 2025 · 0 repositories · arXiv:2503.04971
-
Leveraging Large Language Models to Address Data Scarcity in Machine Learning: Applications in Graphene Synthesis 6 Mar 2025 · 1 repository · arXiv:2503.04870
-
Towards Autonomous Reinforcement Learning for Real-World Robotic Manipulation with Large Language Models 6 Mar 2025 · 0 repositories · arXiv:2503.04280
-
Large language models in finance : what is financial sentiment? 5 Mar 2025 · 0 repositories · arXiv:2503.03612
-
MA-LoT: Multi-Agent Lean-based Long Chain-of-Thought Reasoning enhances Formal Theorem Proving 5 Mar 2025 · 1 repository · arXiv:2503.03205
-
Pretrained LLMs as Real-Time Controllers for Robot Operated Serial Production Line 5 Mar 2025 · 0 repositories · arXiv:2503.03889
-
RiskAgent: Autonomous Medical AI Copilot for Generalist Risk Prediction 5 Mar 2025 · 0 repositories · arXiv:2503.03802
-
CoServe: Efficient Collaboration-of-Experts (CoE) Model Inference with Limited Memory 4 Mar 2025 · 0 repositories · arXiv:2503.02354
-
Weak-to-Strong Generalization Even in Random Feature Networks, Provably 4 Mar 2025 · 0 repositories · arXiv:2503.02877
-
AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification 3 Mar 2025 · 0 repositories · arXiv:2503.01940
-
Unmasking Digital Falsehoods: A Comparative Analysis of LLM-Based Misinformation Detection Strategies 2 Mar 2025 · 0 repositories · arXiv:2503.00724
-
PodAgent: A Comprehensive Framework for Podcast Generation 1 Mar 2025 · 1 repository · arXiv:2503.00455
-
Psychological Counseling Ability of Large Language Models 1 Mar 2025 · 0 repositories · arXiv:2503.07627
-
Learning to Align Multi-Faceted Evaluation: A Unified and Robust Framework 26 Feb 2025 · 0 repositories · arXiv:2502.18874
-
MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering 26 Feb 2025 · 0 repositories · arXiv:2502.18993
-
Reimagining Personal Data: Unlocking the Potential of AI-Generated Images in Personal Data Meaning-Making 26 Feb 2025 · 0 repositories · arXiv:2502.18853
-
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning 26 Feb 2025 · 0 repositories · arXiv:2502.19622
-
Assessing Large Language Models in Agentic Multilingual National Bias 25 Feb 2025 · 0 repositories · arXiv:2502.17945
-
Bayesian Optimization for Controlled Image Editing via LLMs 25 Feb 2025 · 0 repositories · arXiv:2502.18116
-
Stackelberg Game Preference Optimization for Data-Efficient Alignment of Language Models 25 Feb 2025 · 0 repositories · arXiv:2502.18099
-
Are Large Language Models Good Data Preprocessors? 24 Feb 2025 · 0 repositories · arXiv:2502.16790
-
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding) 24 Feb 2025 · 0 repositories · arXiv:2502.17169
-
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation 22 Feb 2025 · 0 repositories · arXiv:2502.16022
-
Uncertainty-Aware Fusion: An Ensemble Framework for Mitigating Hallucinations in Large Language Models 22 Feb 2025 · 0 repositories · arXiv:2503.05757
-
Auto-Bench: An Automated Benchmark for Scientific Discovery in LLMs 21 Feb 2025 · 0 repositories · arXiv:2502.15224
-
AutoMedPrompt: A New Framework for Optimizing LLM Medical Prompts Using Textual Gradients 21 Feb 2025 · 0 repositories · arXiv:2502.15944
-
Comparative Analysis of Large Language Models for Context-Aware Code Completion using SAFIM Framework 21 Feb 2025 · 0 repositories · arXiv:2502.15243
-
Cross-Format Retrieval-Augmented Generation in XR with LLMs for Context-Aware Maintenance Assistance 21 Feb 2025 · 0 repositories · arXiv:2502.15604
-
Single-pass Detection of Jailbreaking Input in Large Language Models 21 Feb 2025 · 0 repositories · arXiv:2502.15435
-
TurboFuzzLLM: Turbocharging Mutation-based Fuzzing for Effectively Jailbreaking Large Language Models in Practice 21 Feb 2025 · 1 repository · arXiv:2502.18504
-
Argument-Based Comparative Question Answering Evaluation Benchmark 20 Feb 2025 · 0 repositories · arXiv:2502.14476
-
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model 20 Feb 2025 · 0 repositories · arXiv:2502.15832
-
Do LLMs Consider Security? An Empirical Study on Responses to Programming Questions 20 Feb 2025 · 0 repositories · arXiv:2502.14202
-
From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT 20 Feb 2025 · 0 repositories · arXiv:2502.14714
-
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning 20 Feb 2025 · 0 repositories · arXiv:2502.14356
-
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding 20 Feb 2025 · 0 repositories · arXiv:2502.14949
-
PaperHelper: Knowledge-Based LLM QA Paper Reading Assistant 20 Feb 2025 · 0 repositories · arXiv:2502.14271
-
Tabular Embeddings for Tables with Bi-Dimensional Hierarchical Metadata and Nesting 20 Feb 2025 · 0 repositories · arXiv:2502.15819
-
Extracting Social Connections from Finnish Karelian Refugee Interviews Using LLMs 19 Feb 2025 · 0 repositories · arXiv:2502.13566
-
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education 19 Feb 2025 · 0 repositories · arXiv:2502.13789
-
STaR-SQL: Self-Taught Reasoner for Text-to-SQL 19 Feb 2025 · 0 repositories · arXiv:2502.13550
-
Language Models are Few-Shot Graders 18 Feb 2025 · 0 repositories · arXiv:2502.13337
-
MatterChat: A Multi-Modal LLM for Material Science 18 Feb 2025 · 0 repositories · arXiv:2502.13107
-
Testing Prompt Engineering Methods for Knowledge Extraction from Text 18 Feb 2025 · 1 repository
-
CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation 17 Feb 2025 · 0 repositories · arXiv:2502.11703
-
MuSC: Improving Complex Instruction Following with Multi-granularity Self-Contrastive Training 17 Feb 2025 · 1 repository · arXiv:2502.11541
-
SmartLLM: Smart Contract Auditing using Custom Generative AI 17 Feb 2025 · 0 repositories · arXiv:2502.13167
-
Empirical evaluation of LLMs in predicting fixes of Configuration bugs in Smart Home System 16 Feb 2025 · 0 repositories · arXiv:2502.10953
-
Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models 16 Feb 2025 · 1 repository · arXiv:2502.11075
-
Performance Review on LLM for solving leetcode problems 16 Feb 2025 · 0 repositories · arXiv:2502.15770
-
Vendi-RAG: Adaptively Trading-Off Diversity And Quality Significantly Improves Retrieval Augmented Generation With LLMs 16 Feb 2025 · 0 repositories · arXiv:2502.11228
-
Can Uniform Meaning Representation Help GPT-4 Translate from Indigenous Languages? 13 Feb 2025 · 0 repositories · arXiv:2502.08900
-
Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs 13 Feb 2025 · 0 repositories · arXiv:2502.09156
-
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? 13 Feb 2025 · 1 repository · arXiv:2502.09416Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance 12 Feb 2025 · 1 repository · arXiv:2502.08127Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
OpenGrok: Enhancing SNS Data Processing with Distilled Knowledge and Mask-like Mechanisms 11 Feb 2025 · 1 repository · arXiv:2502.07312
-
Large Language Models for In-File Vulnerability Localization Can Be "Lost in the End" 9 Feb 2025 · 0 repositories · arXiv:2502.06898
-
Dynamic Noise Preference Optimization for LLM Self-Improvement via Synthetic Data 8 Feb 2025 · 0 repositories · arXiv:2502.05400
-
Can Large Language Models Understand Intermediate Representations? 7 Feb 2025 · 0 repositories · arXiv:2502.06854
-
Unsafe LLM-Based Search: Quantitative Analysis and Mitigation of Safety Risks in AI Web Search 7 Feb 2025 · 0 repositories · arXiv:2502.04951
-
SMI: An Information-Theoretic Metric for Predicting Model Knowledge Solely from Pre-Training Signals 6 Feb 2025 · 1 repository · arXiv:2502.04066
-
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation 6 Feb 2025 · 0 repositories · arXiv:2502.06843
-
OPTIC: Optimizing Patient-Provider Triaging & Improving Communications in Clinical Operations using GPT-4 Data Labeling and Model Distillation 5 Feb 2025 · 0 repositories · arXiv:2503.05701
-
Exploring the Panorama of Anxiety Levels: A Multi-Scenario Study Based on Human-Centric Anxiety Level Detection and Personalized Guidance 4 Feb 2025 · 0 repositories · arXiv:2503.15527
-
LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models 4 Feb 2025 · 1 repository · arXiv:2502.02441
-
Open Foundation Models in Healthcare: Challenges, Paradoxes, and Opportunities with GenAI Driven Personalized Prescription 4 Feb 2025 · 0 repositories · arXiv:2502.04356
-
Toward Neurosymbolic Program Comprehension 3 Feb 2025 · 0 repositories · arXiv:2502.01806
-
Benchmark on Peer Review Toxic Detection: A Challenging Task with a New Dataset 1 Feb 2025 · 0 repositories · arXiv:2502.01676
-
Do LLMs Strategically Reveal, Conceal, and Infer Information? A Theoretical and Empirical Analysis in The Chameleon Game 31 Jan 2025 · 1 repository · arXiv:2501.19398
-
Homogeneity Bias as Differential Sampling Uncertainty in Language Models 31 Jan 2025 · 0 repositories · arXiv:2501.19337
-
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media 31 Jan 2025 · 0 repositories · arXiv:2502.01658
-
Evaluating Large Language Models in Vulnerability Detection Under Variable Context Windows 30 Jan 2025 · 0 repositories · arXiv:2502.00064
-
GENIE: Generative Note Information Extraction model for structuring EHR data 30 Jan 2025 · 0 repositories · arXiv:2501.18435
-
Survey and Improvement Strategies for Gene Prioritization with Large Language Models 30 Jan 2025 · 0 repositories · arXiv:2501.18794
-
Unraveling the Capabilities of Language Models in News Summarization 30 Jan 2025 · 1 repository · arXiv:2501.18128
-
Hybrid Graphs for Table-and-Text based Question Answering using LLMs 29 Jan 2025 · 0 repositories · arXiv:2501.17767
-
Leveraging In-Context Learning and Retrieval-Augmented Generation for Automatic Question Generation in Educational Domains 29 Jan 2025 · 0 repositories · arXiv:2501.17397
-
JRE-L: Journalist, Reader, and Editor LLMs in the Loop for Science Journalism for the General Audience 28 Jan 2025 · 1 repository · arXiv:2501.16865
-
Scenario Understanding of Traffic Scenes Through Large Visual Language Models 28 Jan 2025 · 0 repositories · arXiv:2501.17131
-
LCTG Bench: LLM Controlled Text Generation Benchmark 27 Jan 2025 · 1 repository · arXiv:2501.15875
-
Adapting Biomedical Abstracts into Plain language using Large Language Models 26 Jan 2025 · 0 repositories · arXiv:2501.15700