Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 6
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 6 of 29: papers 501 to 600 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
LLMs: A Game-Changer for Software Engineers? 1 Nov 2024 · 0 repositories · arXiv:2411.00932
-
Self-Evolved Reward Learning for LLMs 1 Nov 2024 · 1 repository · arXiv:2411.00418Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Desert Camels and Oil Sheikhs: Arab-Centric Red Teaming of Frontier LLMs 31 Oct 2024 · 0 repositories · arXiv:2410.24049
-
Large Language Models for Patient Comments Multi-Label Classification 31 Oct 2024 · 0 repositories · arXiv:2410.23528
-
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation 31 Oct 2024 · 1 repository · arXiv:2411.00073Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Danoliteracy of Generative, Large Language Models 30 Oct 2024 · 0 repositories · arXiv:2410.22839
-
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations 30 Oct 2024 · 0 repositories · arXiv:2410.22821Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
SciPIP: An LLM-based Scientific Paper Idea Proposer 30 Oct 2024 · 1 repository · arXiv:2410.23166Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts 29 Oct 2024 · 1 repository · arXiv:2410.22143Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
CFSafety: Comprehensive Fine-grained Safety Assessment for LLMs 29 Oct 2024 · 0 repositories · arXiv:2410.21695
-
Self-Preference Bias in LLM-as-a-Judge 29 Oct 2024 · 0 repositories · arXiv:2410.21819
-
Topic-Conversation Relevance (TCR) Dataset and Benchmarks 29 Oct 2024 · 1 repository · arXiv:2411.00038Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
A Simple Yet Effective Corpus Construction Framework for Indonesian Grammatical Error Correction 28 Oct 2024 · 1 repository · arXiv:2410.20838
-
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models 28 Oct 2024 · 1 repository · arXiv:2410.21195
-
CT2C-QA: Multimodal Question Answering over Chinese Text, Table and Chart 28 Oct 2024 · 0 repositories · arXiv:2410.21414
-
Gender Bias in LLM-generated Interview Responses 28 Oct 2024 · 0 repositories · arXiv:2410.20739
-
Fine-Grained and Multi-Dimensional Metrics for Document-Level Machine Translation 28 Oct 2024 · 1 repository · arXiv:2410.20941
-
Is GPT-4 Less Politically Biased than GPT-3.5? A Renewed Investigation of ChatGPT's Political Biases 28 Oct 2024 · 0 repositories · arXiv:2410.21008
-
SandboxAQ's submission to MRL 2024 Shared Task on Multi-lingual Multi-task Information Retrieval 28 Oct 2024 · 0 repositories · arXiv:2410.21501
-
Malinowski in the Age of AI: Can large language models create a text game based on an anthropological classic? 27 Oct 2024 · 0 repositories · arXiv:2410.20536
-
Sequential Large Language Model-Based Hyper-parameter Optimization 27 Oct 2024 · 1 repository · arXiv:2410.20302
-
Think Carefully and Check Again! Meta-Generation Unlocking LLMs for Low-Resource Cross-Lingual Summarization 26 Oct 2024 · 0 repositories · arXiv:2410.20021
-
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs 25 Oct 2024 · 0 repositories · arXiv:2410.19317
-
GPT-4o System Card 25 Oct 2024 · 0 repositories · arXiv:2410.21276
-
KAHANI: Culturally-Nuanced Visual Storytelling Pipeline for Non-Western Cultures 25 Oct 2024 · 0 repositories · arXiv:2410.19419
-
Robot Behavior Personalization from Sparse User Feedback 25 Oct 2024 · 0 repositories · arXiv:2410.19219
-
Aggregated Knowledge Model: Enhancing Domain-Specific QA with Fine-Tuned and Retrieval-Augmented Generation Models 24 Oct 2024 · 0 repositories · arXiv:2410.18344
-
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning 24 Oct 2024 · 0 repositories · arXiv:2410.18955
-
CAMEL-Bench: A Comprehensive Arabic LMM Benchmark 24 Oct 2024 · 1 repository · arXiv:2410.18976
-
Iterative Self-Tuning LLMs for Enhanced Jailbreaking Capabilities 24 Oct 2024 · 1 repository · arXiv:2410.18469
-
Little Giants: Synthesizing High-Quality Embedding Data at Scale 24 Oct 2024 · 1 repository · arXiv:2410.18634Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
LOGO -- Long cOntext aliGnment via efficient preference Optimization 24 Oct 2024 · 1 repository · arXiv:2410.18533
-
Prompting and Fine-Tuning of Small LLMs for Length-Controllable Telephone Call Summarization 24 Oct 2024 · 0 repositories · arXiv:2410.18624
-
ToolFlow: Boosting LLM Tool-Calling Through Natural and Coherent Dialogue Synthesis 24 Oct 2024 · 0 repositories · arXiv:2410.18447
-
CLR-Bench: Evaluating Large Language Models in College-level Reasoning 23 Oct 2024 · 0 repositories · arXiv:2410.17558
-
From PDFs to Structured Data: Utilizing LLM Analysis in Sports Database Management 23 Oct 2024 · 0 repositories · arXiv:2410.17619
-
Gazelle: An Instruction Dataset for Arabic Writing Assistance 23 Oct 2024 · 0 repositories · arXiv:2410.18163
-
A Statistical Analysis of LLMs' Self-Evaluation Using Proverbs 22 Oct 2024 · 0 repositories · arXiv:2410.16640
-
An Eye for an AI: Evaluating GPT-4o's Visual Perception Skills and Geometric Reasoning Skills Using Computer Graphics Questions 22 Oct 2024 · 0 repositories · arXiv:2410.16991
-
Automated Spinal MRI Labelling from Reports Using a Large Language Model 22 Oct 2024 · 1 repository · arXiv:2410.17235
-
In Context Learning and Reasoning for Symbolic Regression with Large Language Models 22 Oct 2024 · 1 repository · arXiv:2410.17448
-
An Efficient System for Automatic Map Storytelling -- A Case Study on Historical Maps 21 Oct 2024 · 1 repository · arXiv:2410.15780
-
CausalGraph2LLM: Evaluating LLMs for Causal Queries 21 Oct 2024 · 1 repository · arXiv:2410.15939
-
Reflection-Bench: probing AI intelligence with reflection 21 Oct 2024 · 1 repository · arXiv:2410.16270
-
Students Rather Than Experts: A New AI For Education Pipeline To Model More Human-Like And Personalised Early Adolescences 21 Oct 2024 · 0 repositories · arXiv:2410.15701
-
Back to School: Translation Using Grammar Books 20 Oct 2024 · 1 repository · arXiv:2410.15263Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Does ChatGPT Have a Poetic Style? 20 Oct 2024 · 1 repository · arXiv:2410.15299
-
Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations 20 Oct 2024 · 0 repositories · arXiv:2410.15442
-
Training Language Models to Critique With Multi-agent Feedback 20 Oct 2024 · 0 repositories · arXiv:2410.15287
-
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation 19 Oct 2024 · 1 repository · arXiv:2410.14948
-
CausalChat: Interactive Causal Model Development and Refinement Using Large Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14146
-
CELI: Controller-Embedded Language Model Interactions 18 Oct 2024 · 0 repositories · arXiv:2410.14627
-
DFlow: Diverse Dialogue Flow Simulation with Large Language Models 18 Oct 2024 · 0 repositories · arXiv:2410.14853
-
Flame quality monitoring of flare stack based on deep visual features 18 Oct 2024 · 0 repositories · arXiv:2410.19823
-
Good Parenting is all you need -- Multi-agentic LLM Hallucination Mitigation 18 Oct 2024 · 0 repositories · arXiv:2410.14262
-
Novel Development of LLM Driven mCODE Data Model for Improved Clinical Trial Matching to Enable Standardization and Interoperability in Oncology Research 18 Oct 2024 · 0 repositories · arXiv:2410.19826
-
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 18 Oct 2024 · 1 repository · arXiv:2410.14211Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
TimeSeriesExam: A time series understanding exam 18 Oct 2024 · 1 repository · arXiv:2410.14752Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Better to Ask in English: Evaluation of Large Language Models on English, Low-resource and Cross-Lingual Settings 17 Oct 2024 · 0 repositories · arXiv:2410.13153
-
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection 17 Oct 2024 · 0 repositories · arXiv:2410.13464
-
Looking Inward: Language Models Can Learn About Themselves by Introspection 17 Oct 2024 · 1 repository · arXiv:2410.13787Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback 17 Oct 2024 · 1 repository · arXiv:2410.13191
-
Measuring and Modifying the Readability of English Texts with GPT-4 17 Oct 2024 · 1 repository · arXiv:2410.14028
-
SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation 17 Oct 2024 · 1 repository · arXiv:2410.13293
-
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs 17 Oct 2024 · 0 repositories · arXiv:2410.14057Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
AT-RAG: An Adaptive RAG Model Enhancing Query Efficiency with Topic Filtering and Iterative Reasoning 16 Oct 2024 · 1 repository · arXiv:2410.12886
-
CCSBench: Evaluating Compositional Controllability in LLMs for Scientific Document Summarization 16 Oct 2024 · 0 repositories · arXiv:2410.12601
-
Evaluating Morphological Compositional Generalization in Large Language Models 16 Oct 2024 · 1 repository · arXiv:2410.12656Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Identifying Task Groupings for Multi-Task Learning Using Pointwise V-Usable Information 16 Oct 2024 · 0 repositories · arXiv:2410.12774
-
MIRROR: A Novel Approach for the Automated Evaluation of Open-Ended Question Generation 16 Oct 2024 · 0 repositories · arXiv:2410.12893
-
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation 16 Oct 2024 · 1 repository · arXiv:2410.12916Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation 16 Oct 2024 · 0 repositories · arXiv:2410.12222
-
Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Generator-Validator Fine-tuning 16 Oct 2024 · 0 repositories · arXiv:2410.12164
-
Cognitive Overload Attack:Prompt Injection for Long Context 15 Oct 2024 · 1 repository · arXiv:2410.11272
-
De-jargonizing Science for Journalists with GPT-4: A Pilot Study 15 Oct 2024 · 1 repository · arXiv:2410.12069
-
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models 15 Oct 2024 · 1 repository · arXiv:2410.11459
-
Selection-p: Self-Supervised Task-Agnostic Prompt Compression for Faithfulness and Transferability 15 Oct 2024 · 0 repositories · arXiv:2410.11786
-
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings 15 Oct 2024 · 0 repositories · arXiv:2410.12046
-
Code-Mixer Ya Nahi: Novel Approaches to Measuring Multilingual LLMs' Code-Mixing Capabilities 14 Oct 2024 · 0 repositories · arXiv:2410.11079
-
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning 14 Oct 2024 · 0 repositories · arXiv:2410.10735
-
FormalAlign: Automated Alignment Evaluation for Autoformalization 14 Oct 2024 · 1 repository · arXiv:2410.10135Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Gender Bias of LLM in Economics: An Existentialism Perspective 14 Oct 2024 · 0 repositories · arXiv:2410.19775
-
Generative AI and Its Impact on Personalized Intelligent Tutoring Systems 14 Oct 2024 · 0 repositories · arXiv:2410.10650
-
3DS: Decomposed Difficulty Data Selection's Case Study on LLM Medical Domain Adaptation 13 Oct 2024 · 0 repositories · arXiv:2410.10901
-
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs 13 Oct 2024 · 1 repository · arXiv:2410.09775
-
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis 13 Oct 2024 · 0 repositories · arXiv:2410.12867
-
HARDMath: A Benchmark Dataset for Challenging Problems in Applied Mathematics 13 Oct 2024 · 1 repository · arXiv:2410.09988
-
Extended Japanese Commonsense Morality Dataset with Masked Token and Label Enhancement 12 Oct 2024 · 0 repositories · arXiv:2410.09564
-
GPTON: Generative Pre-trained Transformers enhanced with Ontology Narration for accurate annotation of biological data 12 Oct 2024 · 0 repositories · arXiv:2410.10899
-
AttnGCG: Enhancing Jailbreaking Attacks on LLMs with Attention Manipulation 11 Oct 2024 · 1 repository · arXiv:2410.09040
-
Developing a Pragmatic Benchmark for Assessing Korean Legal Language Understanding in Large Language Models 11 Oct 2024 · 1 repository · arXiv:2410.08731
-
Fine-Tuning In-House Large Language Models to Infer Differential Diagnosis from Radiology Reports 11 Oct 2024 · 0 repositories · arXiv:2410.09234
-
Hypothesis-only Biases in Large Language Model-Elicited Natural Language Inference 11 Oct 2024 · 0 repositories · arXiv:2410.08996
-
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework 11 Oct 2024 · 1 repository · arXiv:2410.12855Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Large Language Models for Medical OSCE Assessment: A Novel Approach to Transcript Analysis 11 Oct 2024 · 0 repositories · arXiv:2410.12858
-
SuperCorrect: Supervising and Correcting Language Models with Error-Driven Insights 11 Oct 2024 · 2 repositories · arXiv:2410.09008Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking Agentic Workflow Generation 10 Oct 2024 · 1 repository · arXiv:2410.07869Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Diversity of Thought Elicits Stronger Reasoning Capabilities in Multi-Agent Debate Frameworks 10 Oct 2024 · 0 repositories · arXiv:2410.12853
-
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency 10 Oct 2024 · 0 repositories · arXiv:2410.07563
-
Prompt Engineering a Schizophrenia Chatbot: Utilizing a Multi-Agent Approach for Enhanced Compliance with Prompt Instructions 10 Oct 2024 · 0 repositories · arXiv:2410.12848