Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 7
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 7 of 29: papers 601 to 700 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Rescriber: Smaller-LLM-Powered User-Led Data Minimization for LLM-Based Chatbots 10 Oct 2024 · 0 repositories · arXiv:2410.11876
-
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.08068Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Think Beyond Size: Adaptive Prompting for More Effective Reasoning 10 Oct 2024 · 0 repositories · arXiv:2410.08130
-
Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs) 10 Oct 2024 · 1 repository · arXiv:2410.07507Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.12851Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
AutoFeedback: An LLM-based Framework for Efficient and Accurate API Request Generation 9 Oct 2024 · 0 repositories · arXiv:2410.06943
-
Detecting Bias and Enhancing Diagnostic Accuracy in Large Language Models for Healthcare 9 Oct 2024 · 0 repositories · arXiv:2410.06566
-
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA 9 Oct 2024 · 0 repositories · arXiv:2410.06524
-
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time 9 Oct 2024 · 1 repository · arXiv:2410.06625Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Improving Data Efficiency via Curating LLM-Driven Rating Systems 9 Oct 2024 · 0 repositories · arXiv:2410.10877
-
Investigating Cost-Efficiency of LLM-Generated Training Data for Conversational Semantic Frame Analysis 9 Oct 2024 · 0 repositories · arXiv:2410.06550
-
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints 9 Oct 2024 · 0 repositories · arXiv:2410.06458
-
TextLap: Customizing Language Models for Text-to-Layout Planning 9 Oct 2024 · 1 repository · arXiv:2410.12844
-
TuringQ: Benchmarking AI Comprehension in Theory of Computation 9 Oct 2024 · 1 repository · arXiv:2410.06547
-
Application of NotebookLM, a Large Language Model with Retrieval-Augmented Generation, for Lung Cancer Staging 8 Oct 2024 · 0 repositories · arXiv:2410.10869
-
Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation 8 Oct 2024 · 0 repositories · arXiv:2410.06094
-
SC-Bench: A Large-Scale Dataset for Smart Contract Auditing 8 Oct 2024 · 1 repository · arXiv:2410.06176
-
Contrastive Learning to Improve Retrieval for Real-world Fact Checking 7 Oct 2024 · 0 repositories · arXiv:2410.04657
-
Generating CAD Code with Vision-Language Models for 3D Designs 7 Oct 2024 · 0 repositories · arXiv:2410.05340
-
Diagnosing Robotics Systems Issues with Large Language Models 6 Oct 2024 · 0 repositories · arXiv:2410.09084
-
MindScope: Exploring cognitive biases in large language models through Multi-Agent Systems 6 Oct 2024 · 1 repository · arXiv:2410.04452
-
ProtocoLLM: Automatic Evaluation Framework of LLMs on Domain-Specific Scientific Protocol Formulation Tasks 6 Oct 2024 · 0 repositories · arXiv:2410.04601
-
ECon: On the Detection and Resolution of Evidence Conflicts 5 Oct 2024 · 1 repository · arXiv:2410.04068
-
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation 5 Oct 2024 · 1 repository · arXiv:2410.04002Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition 4 Oct 2024 · 0 repositories · arXiv:2410.03335
-
Beyond Film Subtitles: Is YouTube the Best Approximation of Spoken Vocabulary? 4 Oct 2024 · 1 repository · arXiv:2410.03240
-
Detecting Machine-Generated Long-Form Content with Latent-Space Variables 4 Oct 2024 · 0 repositories · arXiv:2410.03856
-
SAG: Style-Aligned Article Generation via Model Collaboration 4 Oct 2024 · 0 repositories · arXiv:2410.03137
-
Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis 4 Oct 2024 · 0 repositories · arXiv:2410.03908
-
Towards Linguistically-Aware and Language-Independent Tokenization for Large Language Models (LLMs) 4 Oct 2024 · 0 repositories · arXiv:2410.03568
-
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation 3 Oct 2024 · 0 repositories · arXiv:2410.02725
-
Can LLMs Reliably Simulate Human Learner Actions? A Simulation Authoring Framework for Open-Ended Learning Environments 3 Oct 2024 · 1 repository · arXiv:2410.02110
-
CAX: Cellular Automata Accelerated in JAX 3 Oct 2024 · 1 repository · arXiv:2410.02651
-
Coal Mining Question Answering with LLMs 3 Oct 2024 · 0 repositories · arXiv:2410.02959
-
Defining Knowledge: Bridging Epistemology and Large Language Models 3 Oct 2024 · 0 repositories · arXiv:2410.02499
-
Grounding Large Language Models In Embodied Environment With Imperfect World Models 3 Oct 2024 · 0 repositories · arXiv:2410.02742
-
IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models 3 Oct 2024 · 0 repositories · arXiv:2410.02429
-
Training Language Models on Synthetic Edit Sequences Improves Code Synthesis 3 Oct 2024 · 1 repository · arXiv:2410.02749Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses 2 Oct 2024 · 0 repositories · arXiv:2410.01246
-
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester 2 Oct 2024 · 0 repositories · arXiv:2410.01606
-
ET-Plan-Bench: Embodied Task-level Planning Benchmark Towards Spatial-Temporal Cognition with Foundation Models 2 Oct 2024 · 0 repositories · arXiv:2410.14682
-
MARPLE: A Benchmark for Long-Horizon Inference 2 Oct 2024 · 1 repository · arXiv:2410.01926Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Mind Scramble: Unveiling Large Language Model Psychology Via Typoglycemia 2 Oct 2024 · 1 repository · arXiv:2410.01677
-
On The Adaptation of Unlimiformer for Decoder-Only Transformers 2 Oct 2024 · 0 repositories · arXiv:2410.01637
-
Unleashing the Unseen: Harnessing Benign Datasets for Jailbreaking Large Language Models 1 Oct 2024 · 1 repository · arXiv:2410.00451
-
Creative and Context-Aware Translation of East Asian Idioms with GPT-4 1 Oct 2024 · 1 repository · arXiv:2410.00988
-
Decoding Hate: Exploring Language Models' Reactions to Hate Speech 1 Oct 2024 · 0 repositories · arXiv:2410.00775
-
Insight: A Multi-Modal Diagnostic Pipeline using LLMs for Ocular Surface Disease Diagnosis 1 Oct 2024 · 0 repositories · arXiv:2410.00292
-
Language Enhanced Model for Eye (LEME): An Open-Source Ophthalmology-Specific Large Language Model 1 Oct 2024 · 0 repositories · arXiv:2410.03740
-
RATIONALYST: Pre-training Process-Supervision for Improving Reasoning 1 Oct 2024 · 1 repository · arXiv:2410.01044
-
ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer 30 Sep 2024 · 0 repositories · arXiv:2410.00086
-
CliMB: An AI-enabled Partner for Clinical Predictive Modeling 30 Sep 2024 · 1 repository · arXiv:2410.03736Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Exploring Social Media Image Categorization Using Large Models with Different Adaptation Methods: A Case Study on Cultural Nature's Contributions to People 30 Sep 2024 · 0 repositories · arXiv:2410.00275
-
On The Planning Abilities of OpenAI's o1 Models: Feasibility, Optimality, and Generalizability 30 Sep 2024 · 2 repositories · arXiv:2409.19924
-
Can Models Learn Skill Composition from Examples? 29 Sep 2024 · 0 repositories · arXiv:2409.19808
-
GenTel-Safe: A Unified Benchmark and Shielding Framework for Defending Against Prompt Injection Attacks 29 Sep 2024 · 0 repositories · arXiv:2409.19521
-
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models 29 Sep 2024 · 0 repositories · arXiv:2409.19492
-
See then Tell: Enhancing Key Information Extraction with Vision Grounding 29 Sep 2024 · 0 repositories · arXiv:2409.19573
-
Towards AI-Assisted Protocol Analysis in Design Research: Automating Question Labelling with GPT-4 According to Eris’ (2004) Taxonomy 29 Sep 2024 · 1 repository
-
LML-DAP: Language Model Learning a Dataset for Data-Augmented Prediction 27 Sep 2024 · 1 repository · arXiv:2409.18957
-
Not the Silver Bullet: LLM-enhanced Programming Error Messages are Ineffective in Practice 27 Sep 2024 · 0 repositories · arXiv:2409.18661
-
Open-Nav: Exploring Zero-Shot Vision-and-Language Navigation in Continuous Environment with Open-Source LLMs 27 Sep 2024 · 0 repositories · arXiv:2409.18794
-
DARE: Diverse Visual Question Answering with Robustness Evaluation 26 Sep 2024 · 0 repositories · arXiv:2409.18023
-
Just Say What You Want: Only-prompting Self-rewarding Online Preference Optimization 26 Sep 2024 · 0 repositories · arXiv:2409.17534
-
Predicting Anchored Text from Translation Memories for Machine Translation Using Deep Learning Methods 26 Sep 2024 · 0 repositories · arXiv:2409.17939
-
Retrospective Comparative Analysis of Prostate Cancer In-Basket Messages: Responses from Closed-Domain LLM vs. Clinical Teams 26 Sep 2024 · 1 repository · arXiv:2409.18290
-
The application of GPT-4 in grading design university students' assignment and providing feedback: An exploratory study 26 Sep 2024 · 0 repositories · arXiv:2409.17698
-
Beyond Turing Test: Can GPT-4 Sway Experts' Decisions? 25 Sep 2024 · 0 repositories · arXiv:2409.16710
-
CodeInsight: A Curated Dataset of Practical Coding Solutions from Stack Overflow 25 Sep 2024 · 1 repository · arXiv:2409.16819
-
Post-hoc Reward Calibration: A Case Study on Length Bias 25 Sep 2024 · 1 repository · arXiv:2409.17407Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A Comprehensive Evaluation of Large Language Models on Mental Illnesses 24 Sep 2024 · 0 repositories · arXiv:2409.15687
-
AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment 24 Sep 2024 · 0 repositories · arXiv:2409.16022
-
Task-oriented Prompt Enhancement via Script Generation 24 Sep 2024 · 0 repositories · arXiv:2409.16418
-
A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor? 23 Sep 2024 · 0 repositories · arXiv:2409.15277
-
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs 23 Sep 2024 · 1 repository · arXiv:2409.14866Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
MemeCLIP: Leveraging CLIP Representations for Multimodal Meme Classification 23 Sep 2024 · 1 repository · arXiv:2409.14703
-
PALLM: Evaluating and Enhancing PALLiative Care Conversations with Large Language Models 23 Sep 2024 · 1 repository · arXiv:2409.15188
-
Beyond Words: Evaluating Large Language Models in Transportation Planning 22 Sep 2024 · 0 repositories · arXiv:2409.14516
-
Enhancing LLM-based Autonomous Driving Agents to Mitigate Perception Attacks 22 Sep 2024 · 0 repositories · arXiv:2409.14488
-
Evaluating the Quality of Code Comments Generated by Large Language Models for Novice Programmers 22 Sep 2024 · 0 repositories · arXiv:2409.14368
-
Large Model Based Agents: State-of-the-Art, Cooperation Paradigms, Security and Privacy, and Future Trends 22 Sep 2024 · 0 repositories · arXiv:2409.14457
-
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models 21 Sep 2024 · 1 repository · arXiv:2409.13989Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 0 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators 21 Sep 2024 · 1 repository · arXiv:2409.14037
-
Aligning Language Models Using Follow-up Likelihood as Reward Signal 20 Sep 2024 · 1 repository · arXiv:2409.13948
-
Prompting Large Language Models for Supporting the Differential Diagnosis of Anemia 20 Sep 2024 · 0 repositories · arXiv:2409.15377
-
ShizishanGPT: An Agricultural Large Language Model Integrating Tools and Resources 20 Sep 2024 · 1 repository · arXiv:2409.13537
-
STOP! Benchmarking Large Language Models with Sensitivity Testing on Offensive Progressions 20 Sep 2024 · 1 repository · arXiv:2409.13843
-
Enhancing TinyBERT for Financial Sentiment Analysis Using GPT-Augmented FinBERT Distillation 19 Sep 2024 · 1 repository · arXiv:2409.18999
-
Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering 19 Sep 2024 · 1 repository · arXiv:2409.12784
-
Prompts Are Programs Too! Understanding How Developers Build Software Containing Prompts 19 Sep 2024 · 0 repositories · arXiv:2409.12447
-
TACO-RL: Task Aware Prompt Compression Optimization with Reinforcement Learning 19 Sep 2024 · 0 repositories · arXiv:2409.13035
-
What Would You Ask When You First Saw a²+b²=c²? Evaluating LLM on Curiosity-Driven Questioning 19 Sep 2024 · 0 repositories · arXiv:2409.17172
-
From Lists to Emojis: How Format Bias Affects Model Alignment 18 Sep 2024 · 0 repositories · arXiv:2409.11704
-
Harnessing LLMs for API Interactions: A Framework for Classification and Synthetic Data Generation 18 Sep 2024 · 0 repositories · arXiv:2409.11703
-
Investigating Context-Faithfulness in Large Language Models: The Roles of Memory Strength and Evidence Style 17 Sep 2024 · 0 repositories · arXiv:2409.10955
-
Sparks of Artificial General Intelligence(AGI) in Semiconductor Material Science: Early Explorations into the Next Frontier of Generative AI-Assisted Electron Micrograph Analysis 17 Sep 2024 · 0 repositories · arXiv:2409.12244
-
GPT takes the SAT: Tracing changes in Test Difficulty and Math Performance of Students 16 Sep 2024 · 0 repositories · arXiv:2409.10750
-
LLMs for clinical risk prediction 16 Sep 2024 · 0 repositories · arXiv:2409.10191
-
MindGuard: Towards Accessible and Sitgma-free Mental Health First Aid via Edge LLM 16 Sep 2024 · 0 repositories · arXiv:2409.10064
-
SelECT-SQL: Self-correcting ensemble Chain-of-Thought for Text-to-SQL 16 Sep 2024 · 1 repository · arXiv:2409.10007