Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 16
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 16 of 29: papers 1,501 to 1,600 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution 26 Mar 2024 · 0 repositories · arXiv:2403.17927
-
Supervisory Prompt Training 26 Mar 2024 · 0 repositories · arXiv:2403.18051
-
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs 26 Mar 2024 · 0 repositories · arXiv:2403.17856
-
A comparison of Human, GPT-3.5, and GPT-4 Performance in a University-Level Coding Course 25 Mar 2024 · 1 repository · arXiv:2403.16977
-
Do LLM Agents Have Regret? A Case Study in Online Learning and Games 25 Mar 2024 · 0 repositories · arXiv:2403.16843
-
Text Understanding in GPT-4 vs Humans 25 Mar 2024 · 0 repositories · arXiv:2403.17196
-
Linear Cross-document Event Coreference Resolution with X-AMR 25 Mar 2024 · 1 repository · arXiv:2404.08656
-
SeSaMe: A Framework to Simulate Self-Reported Ground Truth for Mental Health Sensing Studies 25 Mar 2024 · 1 repository · arXiv:2403.17219
-
State Space Models as Foundation Models: A Control Theoretic Overview 25 Mar 2024 · 1 repository · arXiv:2403.16899Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 7 pointer-only (licence)
-
EAGLE: A Domain Generalization Framework for AI-generated Text Detection 23 Mar 2024 · 0 repositories · arXiv:2403.15690
-
LlamBERT: Large-scale low-cost data annotation in NLP 23 Mar 2024 · 1 repository · arXiv:2403.15938
-
Using Large Language Models for OntoClean-based Ontology Refinement 23 Mar 2024 · 0 repositories · arXiv:2403.15864
-
Can large language models explore in-context? 22 Mar 2024 · 0 repositories · arXiv:2403.15371
-
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation 22 Mar 2024 · 1 repository · arXiv:2403.14965
-
Construction of a Japanese Financial Benchmark for Large Language Models 22 Mar 2024 · 1 repository · arXiv:2403.15062Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
ESG Classification by Implicit Rule Learning via GPT-4 22 Mar 2024 · 0 repositories · arXiv:2403.15040
-
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization 22 Mar 2024 · 0 repositories · arXiv:2403.15330
-
A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students' Formative Assessment Responses in Science 21 Mar 2024 · 0 repositories · arXiv:2403.14565
-
Assessing the Utility of Large Language Models for Phenotype-Driven Gene Prioritization in Rare Genetic Disorder Diagnosis 21 Mar 2024 · 0 repositories · arXiv:2403.14801
-
Exploring the Potential of Large Language Models in Graph Generation 21 Mar 2024 · 0 repositories · arXiv:2403.14358
-
K-Act2Emo: Korean Commonsense Knowledge Graph for Indirect Emotional Expression 21 Mar 2024 · 1 repository · arXiv:2403.14253
-
LLM-based Extraction of Contradictions from Patents 21 Mar 2024 · 0 repositories · arXiv:2403.14258
-
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy 21 Mar 2024 · 0 repositories · arXiv:2403.14589
-
Facilitating Pornographic Text Detection for Open-Domain Dialogue Systems via Knowledge Distillation of Large Language Models 20 Mar 2024 · 1 repository · arXiv:2403.13250
-
Natural Language as Policies: Reasoning for Coordinate-Level Embodied Control with LLMs 20 Mar 2024 · 0 repositories · arXiv:2403.13801
-
Automated Data Curation for Robust Language Model Fine-Tuning 19 Mar 2024 · 0 repositories · arXiv:2403.12776
-
Automatic Information Extraction From Employment Tribunal Judgements Using Large Language Models 19 Mar 2024 · 0 repositories · arXiv:2403.12936
-
Efficient Encoder-Decoder Transformer Decoding for Decomposable Tasks 19 Mar 2024 · 1 repository · arXiv:2403.13112
-
End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations 19 Mar 2024 · 1 repository · arXiv:2403.12451Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
LHMKE: A Large-scale Holistic Multi-subject Knowledge Evaluation Benchmark for Chinese Large Language Models 19 Mar 2024 · 0 repositories · arXiv:2403.12601
-
Pragmatic Competence Evaluation of Large Language Models for the Korean Language 19 Mar 2024 · 1 repository · arXiv:2403.12675
-
RankPrompt: Step-by-Step Comparisons Make Language Models Better Reasoners 19 Mar 2024 · 0 repositories · arXiv:2403.12373
-
VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning 19 Mar 2024 · 1 repository · arXiv:2403.13164Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Counting-Stars: A Multi-evidence, Position-aware, and Scalable Benchmark for Evaluating Long-Context Large Language Models 18 Mar 2024 · 1 repository · arXiv:2403.11802
-
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models 18 Mar 2024 · 1 repository · arXiv:2403.12171Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing Taiwanese Hokkien Dual Translation by Exploring and Standardizing of Four Writing Systems 18 Mar 2024 · 1 repository · arXiv:2403.12024
-
Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models 18 Mar 2024 · 1 repository · arXiv:2403.11838
-
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents 18 Mar 2024 · 0 repositories · arXiv:2403.12014
-
GPT-4 as Evaluator: Evaluating Large Language Models on Pest Management in Agriculture 18 Mar 2024 · 0 repositories · arXiv:2403.11858
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments 18 Mar 2024 · 1 repository · arXiv:2403.11807Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 2 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Leveraging Large Language Models to Detect npm Malicious Packages 18 Mar 2024 · 0 repositories · arXiv:2403.12196
-
Aligning Uncertainty: Leveraging LLMs to Analyze Uncertainty Transfer in Text Summarization 17 Mar 2024 · 0 repositories
-
Correcting misinformation on social media with a large language model 17 Mar 2024 · 1 repository · arXiv:2403.11169
-
HumSum: A Personalized Lecture Summarization Tool for Humanities Students Using LLMs 17 Mar 2024 · 0 repositories
-
Can Large Language Models Solve Robot Routing? 16 Mar 2024 · 1 repository · arXiv:2403.10795
-
Large language model-powered chatbots for internationalizing student support in higher education 16 Mar 2024 · 0 repositories · arXiv:2403.14702
-
A Thorough Comparison of Cross-Encoders and LLMs for Reranking SPLADE 15 Mar 2024 · 0 repositories · arXiv:2403.10407
-
A Continued Pretrained LLM Approach for Automatic Medical Note Generation 14 Mar 2024 · 0 repositories · arXiv:2403.09057
-
AraTrust: An Evaluation of Trustworthiness for LLMs in Arabic 14 Mar 2024 · 0 repositories · arXiv:2403.09017
-
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences 14 Mar 2024 · 2 repositories · arXiv:2403.09032Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Evaluating LLMs for Gender Disparities in Notable Persons 14 Mar 2024 · 0 repositories · arXiv:2403.09148
-
LAMP: A Language Model on the Map 14 Mar 2024 · 1 repository · arXiv:2403.09059
-
Sabiá-2: A New Generation of Portuguese Large Language Models 14 Mar 2024 · 0 repositories · arXiv:2403.09887
-
VisionGPT-3D: A Generalized Multimodal Agent for Enhanced 3D Vision Understanding 14 Mar 2024 · 0 repositories · arXiv:2403.09530
-
Can Large Language Models Identify Authorship? 13 Mar 2024 · 1 repository · arXiv:2403.08213
-
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study 13 Mar 2024 · 2 repositories · arXiv:2403.08604Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
Distilling Named Entity Recognition Models for Endangered Species from Large Language Models 13 Mar 2024 · 0 repositories · arXiv:2403.15430
-
Evaluating the Application of Large Language Models to Generate Feedback in Programming Education 13 Mar 2024 · 0 repositories · arXiv:2403.09744
-
Large Language Models are Contrastive Reasoners 13 Mar 2024 · 1 repository · arXiv:2403.08211
-
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion 12 Mar 2024 · 2 repositories · arXiv:2403.07865Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions 12 Mar 2024 · 1 repository · arXiv:2403.07678
-
Rethinking Generative Large Language Model Evaluation for Semantic Comprehension 12 Mar 2024 · 0 repositories · arXiv:2403.07872
-
SIFiD: Reassess Summary Factual Inconsistency Detection with LLM 12 Mar 2024 · 0 repositories · arXiv:2403.07557
-
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models 12 Mar 2024 · 4 repositories · arXiv:2403.07714Syntology official (archive's flag): 3 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 5 pointer-only (licence)
-
Stress index strategy enhanced with financial news sentiment analysis for the equity markets 12 Mar 2024 · 0 repositories · arXiv:2404.00012
-
Towards a clinically accessible radiology foundation model: open-access and lightweight, with automated evaluation 12 Mar 2024 · 1 repository · arXiv:2403.08002Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Amharic LLaMA and LLaVA: Multimodal LLMs for Low Resource Languages 11 Mar 2024 · 1 repository · arXiv:2403.06354
-
Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds 11 Mar 2024 · 0 repositories · arXiv:2403.06609
-
SMART: Automatically Scaling Down Language Models with Accuracy Guarantees for Reduced Processing Fees 11 Mar 2024 · 1 repository · arXiv:2403.13835
-
Unraveling the Mystery of Scaling Laws: Part I 11 Mar 2024 · 0 repositories · arXiv:2403.06563
-
AutoEval Done Right: Using Synthetic Data for Model Evaluation 9 Mar 2024 · 1 repository · arXiv:2403.07008Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ClinicalMamba: A Generative Clinical Language Model on Longitudinal Clinical Notes 9 Mar 2024 · 1 repository · arXiv:2403.05795Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
A Dataset and Benchmark for Hospital Course Summarization with Adapted Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05720
-
An In-depth Evaluation of GPT-4 in Sentence Simplification with Error-based Human Assessment 8 Mar 2024 · 0 repositories · arXiv:2403.04963
-
Are Large Language Models Aligned with People's Social Intuitions for Human-Robot Interactions? 8 Mar 2024 · 1 repository · arXiv:2403.05701
-
Can't Remember Details in Long Documents? You Need Some R&R 8 Mar 2024 · 1 repository · arXiv:2403.05004Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Cost-Performance Optimization for Processing Low-Resource Language Tasks Using Commercial LLMs 8 Mar 2024 · 0 repositories · arXiv:2403.05434
-
How Well Do Multi-modal LLMs Interpret CT Scans? An Auto-Evaluation Framework for Analyses 8 Mar 2024 · 0 repositories · arXiv:2403.05680
-
ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05266Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 8 Mar 2024 · 1 repository · arXiv:2403.05530
-
LLM4Decompile: Decompiling Binary Code with Large Language Models 8 Mar 2024 · 1 repository · arXiv:2403.05286Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
RAT: Retrieval Augmented Thoughts Elicit Context-Aware Reasoning in Long-Horizon Generation 8 Mar 2024 · 1 repository · arXiv:2403.05313
-
Will GPT-4 Run DOOM? 8 Mar 2024 · 0 repositories · arXiv:2403.05468
-
Feedback-Generation for Programming Exercises With GPT-4 7 Mar 2024 · 0 repositories · arXiv:2403.04449
-
HaluEval-Wild: Evaluating Hallucinations of Language Models in the Wild 7 Mar 2024 · 1 repository · arXiv:2403.04307
-
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error 7 Mar 2024 · 1 repository · arXiv:2403.04746Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
Assessing the Aesthetic Evaluation Capabilities of GPT-4 with Vision: Insights from Group and Individual Assessments 6 Mar 2024 · 0 repositories · arXiv:2403.03594
-
Can Large Language Models do Analytical Reasoning? 6 Mar 2024 · 0 repositories · arXiv:2403.04031
-
Designing Informative Metrics for Few-Shot Example Selection 6 Mar 2024 · 0 repositories · arXiv:2403.03861
-
General2Specialized LLMs Translation for E-commerce 6 Mar 2024 · 0 repositories · arXiv:2403.03689
-
PPTC-R benchmark: Towards Evaluating the Robustness of Large Language Models for PowerPoint Task Completion 6 Mar 2024 · 1 repository · arXiv:2403.03788
-
Rapidly Developing High-quality Instruction Data and Evaluation Benchmark for Large Language Models with Minimal Human Effort: A Case Study on Japanese 6 Mar 2024 · 2 repositories · arXiv:2403.03690
-
AI Insights: A Case Study on Utilizing ChatGPT Intelligence for Research Paper Analysis 5 Mar 2024 · 0 repositories · arXiv:2403.03293
-
An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4 5 Mar 2024 · 1 repository · arXiv:2403.02839Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments 5 Mar 2024 · 0 repositories · arXiv:2403.03203
-
Emerging Synergies Between Large Language Models and Machine Learning in Ecommerce Recommendations 5 Mar 2024 · 0 repositories · arXiv:2403.02760
-
InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents 5 Mar 2024 · 2 repositories · arXiv:2403.02691Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset 5 Mar 2024 · 1 repository · arXiv:2403.03167
-
Scope of Large Language Models for Mining Emerging Opinions in Online Health Discourse 5 Mar 2024 · 0 repositories · arXiv:2403.03336
-
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection 5 Mar 2024 · 0 repositories · arXiv:2403.03170