Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 8
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 8 of 29: papers 701 to 800 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Detection Made Easy: Potentials of Large Language Models for Solidity Vulnerabilities 15 Sep 2024 · 0 repositories · arXiv:2409.10574
-
GP-GPT: Large Language Model for Gene-Phenotype Mapping 15 Sep 2024 · 0 repositories · arXiv:2409.09825
-
Leveraging Open-Source Large Language Models for Native Language Identification 15 Sep 2024 · 0 repositories · arXiv:2409.09659
-
Unveiling Gender Bias in Large Language Models: Using Teacher's Evaluation in Higher Education As an Example 15 Sep 2024 · 1 repository · arXiv:2409.09652
-
An empirical evaluation of using ChatGPT to summarize disputes for recommending similar labor and employment cases in Chinese 14 Sep 2024 · 0 repositories · arXiv:2409.09280
-
Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI 14 Sep 2024 · 0 repositories · arXiv:2409.09467
-
A RAG Approach for Generating Competency Questions in Ontology Engineering 13 Sep 2024 · 0 repositories · arXiv:2409.08820
-
ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning 13 Sep 2024 · 1 repository · arXiv:2409.08582Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
KodeXv0.1: A Family of State-of-the-Art Financial Large Language Models 13 Sep 2024 · 0 repositories · arXiv:2409.13749
-
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots 12 Sep 2024 · 0 repositories · arXiv:2409.07823
-
How Effectively Do LLMs Extract Feature-Sentiment Pairs from App Reviews? 11 Sep 2024 · 1 repository · arXiv:2409.07162
-
Can We Count on LLMs? The Fixed-Effect Fallacy and Claims of GPT-4 Capabilities 11 Sep 2024 · 0 repositories · arXiv:2409.07638
-
Multimodal Emotion Recognition with Vision-language Prompting and Modality Dropout 11 Sep 2024 · 0 repositories · arXiv:2409.07078
-
SimulBench: Evaluating Language Models with Creative Simulation Tasks 11 Sep 2024 · 0 repositories · arXiv:2409.07641
-
Mapping Biomedical Ontology Terms to IDs: Effect of Domain Prevalence on Prediction Accuracy 11 Sep 2024 · 0 repositories · arXiv:2409.13746
-
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task 10 Sep 2024 · 0 repositories · arXiv:2409.06883
-
Can Large Language Models Unlock Novel Scientific Research Ideas? 10 Sep 2024 · 1 repository · arXiv:2409.06185Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples)
-
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering 10 Sep 2024 · 1 repository · arXiv:2409.06595
-
What is the Role of Small Models in the LLM Era: A Survey 10 Sep 2024 · 1 repository · arXiv:2409.06857
-
Classification performance and reproducibility of GPT-4 omni for information extraction from veterinary electronic health records 9 Sep 2024 · 1 repository · arXiv:2409.13727
-
FairHome: A Fair Housing and Fair Lending Dataset 9 Sep 2024 · 0 repositories · arXiv:2409.05990
-
Identifying the sources of ideological bias in GPT models through linguistic variation in output 9 Sep 2024 · 0 repositories · arXiv:2409.06043
-
Towards Building a Robust Knowledge Intensive Question Answering Model with Large Language Models 9 Sep 2024 · 0 repositories · arXiv:2409.05385
-
VidLPRO: A Video-Language Pre-training Framework for Robotic and Laparoscopic Surgery 7 Sep 2024 · 0 repositories · arXiv:2409.04732
-
AnyMatch -- Efficient Zero-Shot Entity Matching with a Small Language Model 6 Sep 2024 · 1 repository · arXiv:2409.04073
-
Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering 6 Sep 2024 · 0 repositories · arXiv:2409.04181
-
Retrieval Augmented Generation-Based Incident Resolution Recommendation System for IT Support 6 Sep 2024 · 0 repositories · arXiv:2409.13707
-
UI-JEPA: Towards Active Perception of User Intent through Onscreen User Activity 6 Sep 2024 · 0 repositories · arXiv:2409.04081
-
CACER: Clinical Concept Annotations for Cancer Events and Relations 5 Sep 2024 · 1 repository · arXiv:2409.03905
-
MaterialBENCH: Evaluating College-Level Materials Science Problem-Solving Abilities of Large Language Models 5 Sep 2024 · 0 repositories · arXiv:2409.03161
-
xLAM: A Family of Large Action Models to Empower AI Agent Systems 5 Sep 2024 · 1 repository · arXiv:2409.03215
-
Detecting Calls to Action in Multimodal Content: Analysis of the 2021 German Federal Election Campaign on Instagram 4 Sep 2024 · 0 repositories · arXiv:2409.02690
-
How Privacy-Savvy Are Large Language Models? A Case Study on Compliance and Privacy Technical Review 4 Sep 2024 · 0 repositories · arXiv:2409.02375
-
Hypothesizing Missing Causal Variables with LLMs 4 Sep 2024 · 1 repository · arXiv:2409.02604
-
Irrelevant Alternatives Bias Large Language Model Hiring Decisions 4 Sep 2024 · 0 repositories · arXiv:2409.15299
-
Leveraging Large Language Models for Solving Rare MIP Challenges 3 Sep 2024 · 0 repositories · arXiv:2409.04464
-
Self-Instructed Derived Prompt Generation Meets In-Context Learning: Unlocking New Potential of Black-Box LLMs 3 Sep 2024 · 0 repositories · arXiv:2409.01552
-
Large Language Models versus Classical Machine Learning: Performance in COVID-19 Mortality Prediction Using High-Dimensional Tabular Data 2 Sep 2024 · 2 repositories · arXiv:2409.02136
-
Self-Judge: Selective Instruction Following with Alignment Self-Evaluation 2 Sep 2024 · 1 repository · arXiv:2409.00935Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
ToolACE: Winning the Points of LLM Function Calling 2 Sep 2024 · 0 repositories · arXiv:2409.00920
-
An Empirical Study on Information Extraction using Large Language Models 31 Aug 2024 · 0 repositories · arXiv:2409.00369
-
Chatting Up Attachment: Using LLMs to Predict Adult Bonds 31 Aug 2024 · 0 repositories · arXiv:2409.00347
-
Evaluating the Performance of Large Language Models in Competitive Programming: A Multi-Year, Multi-Grade Analysis 31 Aug 2024 · 0 repositories · arXiv:2409.09054
-
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models 31 Aug 2024 · 1 repository · arXiv:2409.00509Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Enhancing Event Reasoning in Large Language Models through Instruction Fine-Tuning with Semantic Causal Graphs 30 Aug 2024 · 0 repositories · arXiv:2409.00209
-
From Text to Emotion: Unveiling the Emotion Annotation Capabilities of LLMs 30 Aug 2024 · 1 repository · arXiv:2408.17026
-
Leveraging a Cognitive Model to Measure Subjective Similarity of Human and GPT-4 Written Content 30 Aug 2024 · 0 repositories · arXiv:2409.00269
-
OrthoDoc: Multimodal Large Language Model for Assisting Diagnosis in Computed Tomography 30 Aug 2024 · 0 repositories · arXiv:2409.09052
-
ProGRes: Prompted Generative Rescoring on ASR n-Best 30 Aug 2024 · 1 repository · arXiv:2409.00217
-
REFFLY: Melody-Constrained Lyrics Editing Model 30 Aug 2024 · 0 repositories · arXiv:2409.00292
-
Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management 29 Aug 2024 · 0 repositories · arXiv:2409.00128
-
Enhancing AI-Driven Psychological Consultation: Layered Prompts with Large Language Models 29 Aug 2024 · 0 repositories · arXiv:2408.16276
-
Enhancing Dialogue Generation in Werewolf Game Through Situation Analysis and Persuasion Strategies 29 Aug 2024 · 0 repositories · arXiv:2408.16586
-
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action 29 Aug 2024 · 1 repository · arXiv:2409.00138Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation 29 Aug 2024 · 0 repositories · arXiv:2408.16730
-
AeroVerse: UAV-Agent Benchmark Suite for Simulating, Pre-training, Finetuning, and Evaluating Aerospace Embodied World Models 28 Aug 2024 · 0 repositories · arXiv:2408.15511
-
ANVIL: Anomaly-based Vulnerability Identification without Labelled Training Data 28 Aug 2024 · 0 repositories · arXiv:2408.16028
-
Evaluating Computational Representations of Character: An Austen Character Similarity Benchmark 28 Aug 2024 · 0 repositories · arXiv:2408.16131
-
Evaluating Named Entity Recognition Using Few-Shot Prompting with Large Language Models 28 Aug 2024 · 1 repository · arXiv:2408.15796
-
FRACTURED-SORRY-Bench: Framework for Revealing Attacks in Conversational Turns Undermining Refusal Efficacy and Defenses over SORRY-Bench (Automated Multi-shot Jailbreaks) 28 Aug 2024 · 0 repositories · arXiv:2408.16163
-
Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions 28 Aug 2024 · 1 repository · arXiv:2408.15787Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Towards Logically Sound Natural Language Reasoning with Logic-Enhanced Language Model Agents 28 Aug 2024 · 1 repository · arXiv:2408.16081
-
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation 28 Aug 2024 · 1 repository · arXiv:2408.15876
-
WebPilot: A Versatile and Autonomous Multi-Agent System for Web Task Execution with Strategic Exploration 28 Aug 2024 · 0 repositories · arXiv:2408.15978
-
Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis 27 Aug 2024 · 0 repositories · arXiv:2409.07463
-
Negation Blindness in Large Language Models: Unveiling the NO Syndrome in Image Generation 27 Aug 2024 · 0 repositories · arXiv:2409.00105
-
Parameter-Efficient Quantized Mixture-of-Experts Meets Vision-Language Instruction Tuning for Semiconductor Electron Micrograph Analysis 27 Aug 2024 · 0 repositories · arXiv:2408.15305
-
Strategic Optimization and Challenges of Large Language Models in Object-Oriented Programming 27 Aug 2024 · 0 repositories · arXiv:2408.14834
-
The Mamba in the Llama: Distilling and Accelerating Hybrid Models 27 Aug 2024 · 2 repositories · arXiv:2408.15237Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 14 harvested samples) · 5 pointer-only (licence)
-
TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering 27 Aug 2024 · 1 repository · arXiv:2408.15299
-
Probing Causality Manipulation of Large Language Models 26 Aug 2024 · 0 repositories · arXiv:2408.14380
-
Prompt-Matcher: Leveraging Large Models to Reduce Uncertainty in Schema Matching Results 24 Aug 2024 · 0 repositories · arXiv:2408.14507
-
IQA-EVAL: Automatic Evaluation of Human-Model Interactive Question Answering 24 Aug 2024 · 0 repositories · arXiv:2408.13545
-
Utilizing Large Language Models for Named Entity Recognition in Traditional Chinese Medicine against COVID-19 Literature: Comparative Study 24 Aug 2024 · 0 repositories · arXiv:2408.13501
-
Knowledge Graph Modeling-Driven Large Language Model Operating System (LLM OS) for Task Automation in Process Engineering Problem-Solving 23 Aug 2024 · 0 repositories · arXiv:2408.14494
-
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates 23 Aug 2024 · 1 repository · arXiv:2408.13006Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Can LLMs Understand Social Norms in Autonomous Driving Games? 22 Aug 2024 · 0 repositories · arXiv:2408.12680
-
Enhancing Automated Program Repair with Solution Design 22 Aug 2024 · 0 repositories · arXiv:2408.12056
-
Optimizing Performance: How Compact Models Match or Exceed GPT's Classification Capabilities through Fine-Tuning 22 Aug 2024 · 0 repositories · arXiv:2409.11408
-
RuleAlign: Making Large Language Models Better Physicians with Diagnostic Rule Alignment 22 Aug 2024 · 0 repositories · arXiv:2408.12579
-
Towards Evaluating and Building Versatile Large Language Models for Medicine 22 Aug 2024 · 1 repository · arXiv:2408.12547Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
BURExtract-Llama: An LLM for Clinical Concept Extraction in Breast Ultrasound Reports 21 Aug 2024 · 0 repositories · arXiv:2408.11334
-
D-RMGPT: Robot-assisted collaborative tasks driven by large multimodal models 21 Aug 2024 · 0 repositories · arXiv:2408.11761
-
DocTabQA: Answering Questions from Long Documents Using Tables 21 Aug 2024 · 1 repository · arXiv:2408.11490
-
Exploring Large Language Models for Feature Selection: A Data-centric Perspective 21 Aug 2024 · 0 repositories · arXiv:2408.12025
-
Leveraging Fine-Tuned Retrieval-Augmented Generation with Long-Context Support: For 3GPP Standards 21 Aug 2024 · 1 repository · arXiv:2408.11775
-
SarcasmBench: Towards Evaluating Large Language Models on Sarcasm Understanding 21 Aug 2024 · 0 repositories · arXiv:2408.11319
-
Crafting Tomorrow's Headlines: Neural News Generation and Detection in English, Turkish, Hungarian, and Persian 20 Aug 2024 · 0 repositories · arXiv:2408.10724
-
Dr.Academy: A Benchmark for Evaluating Questioning Capability in Education for Large Language Models 20 Aug 2024 · 0 repositories · arXiv:2408.10947
-
How Well Do Large Language Models Serve as End-to-End Secure Code Agents for Python? 20 Aug 2024 · 0 repositories · arXiv:2408.10495
-
Open-FinLLMs: Open Multimodal Large Language Models for Financial Applications 20 Aug 2024 · 0 repositories · arXiv:2408.11878
-
Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation 20 Aug 2024 · 1 repository · arXiv:2408.11053Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Soda-Eval: Open-Domain Dialogue Evaluation in the age of LLMs 20 Aug 2024 · 1 repository · arXiv:2408.10902
-
Large Language Models for Classical Chinese Poetry Translation: Benchmarking, Evaluating, and Improving 19 Aug 2024 · 0 repositories · arXiv:2408.09945
-
Edge-Cloud Collaborative Motion Planning for Autonomous Driving with Large Language Models 19 Aug 2024 · 0 repositories · arXiv:2408.09972
-
Self-Directed Turing Test for Large Language Models 19 Aug 2024 · 0 repositories · arXiv:2408.09853
-
Out-of-distribution generalization via composition: a lens through induction heads in Transformers 18 Aug 2024 · 1 repository · arXiv:2408.09503Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Sentiment analysis of preservice teachers' reflections using a large language model 17 Aug 2024 · 0 repositories · arXiv:2408.11862
-
TableBench: A Comprehensive and Complex Benchmark for Table Question Answering 17 Aug 2024 · 0 repositories · arXiv:2408.09174
-
Unraveling Text Generation in LLMs: A Stochastic Differential Equation Approach 17 Aug 2024 · 0 repositories · arXiv:2408.11863