Methods › Natural Language Processing › Language Models › GPT-4 › Papers, page 9
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Page 9 of 29: papers 801 to 900 of 2,870, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Blockchain-Enabled Accountability in Data Supply Chain: A Data Bill of Materials Approach 16 Aug 2024 · 0 repositories · arXiv:2408.08536
-
Can Large Language Models Improve the Adversarial Robustness of Graph Neural Networks? 16 Aug 2024 · 1 repository · arXiv:2408.08685
-
Persona is a Double-edged Sword: Mitigating the Negative Impact of Role-playing Prompts in Zero-shot Reasoning Tasks 16 Aug 2024 · 0 repositories · arXiv:2408.08631
-
See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses 16 Aug 2024 · 1 repository · arXiv:2408.08978
-
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models 15 Aug 2024 · 1 repository · arXiv:2408.07983
-
Benchmarking the Capabilities of Large Language Models in Transportation System Engineering: Accuracy, Consistency, and Reasoning Behaviors 15 Aug 2024 · 0 repositories · arXiv:2408.08302
-
Evaluating the Validity of Word-level Adversarial Attacks with Large Language Models 15 Aug 2024 · 1 repository
-
Leveraging Web-Crawled Data for High-Quality Fine-Tuning 15 Aug 2024 · 1 repository · arXiv:2408.08003
-
MAG-SQL: Multi-Agent Generative Approach with Soft Schema Linking and Iterative Sub-SQL Refinement for Text-to-SQL 15 Aug 2024 · 1 repository · arXiv:2408.07930Syntology official (archive's flag): 21 ran · 21 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 1 honoured, 3 violated, 15 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 25 harvested samples) · 4 pointer-only (licence)
-
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models 15 Aug 2024 · 0 repositories · arXiv:2408.07975
-
CodeMirage: Hallucinations in Code Generated by Large Language Models 14 Aug 2024 · 0 repositories · arXiv:2408.08333
-
A Perspective on Large Language Models, Intelligent Machines, and Knowledge Acquisition 13 Aug 2024 · 0 repositories · arXiv:2408.06598
-
Generative AI for automatic topic labelling 13 Aug 2024 · 0 repositories · arXiv:2408.07003
-
Harnessing Earnings Reports for Stock Predictions: A QLoRA-Enhanced LLM Approach 13 Aug 2024 · 0 repositories · arXiv:2408.06634
-
Leveraging Language Models for Emotion and Behavior Analysis in Education 13 Aug 2024 · 0 repositories · arXiv:2408.06874
-
Using Advanced LLMs to Enhance Smaller LLMs: An Interpretable Knowledge Distillation Approach 13 Aug 2024 · 0 repositories · arXiv:2408.07238
-
Cross-Lingual Conversational Speech Summarization with Large Language Models 12 Aug 2024 · 0 repositories · arXiv:2408.06484
-
Med42-v2: A Suite of Clinical LLMs 12 Aug 2024 · 0 repositories · arXiv:2408.06142
-
The Language of Trauma: Modeling Traumatic Event Descriptions Across Domains with Explainable AI 12 Aug 2024 · 0 repositories · arXiv:2408.05977
-
GPT-4 Emulates Average-Human Emotional Cognition from a Third-Person Perspective 11 Aug 2024 · 0 repositories · arXiv:2408.13718
-
Chain of Condition: Construct, Verify and Solve Conditions for Conditional Question Answering 10 Aug 2024 · 0 repositories · arXiv:2408.05442
-
ChatGPT Meets Iris Biometrics 9 Aug 2024 · 0 repositories · arXiv:2408.04868
-
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners 9 Aug 2024 · 0 repositories · arXiv:2408.05204
-
From Text to Insight: Leveraging Large Language Models for Performance Evaluation in Management 9 Aug 2024 · 0 repositories · arXiv:2408.05328
-
Large Language Models and Thematic Analysis: Human-AI Synergy in Researching Hate Speech on Social Media 9 Aug 2024 · 0 repositories · arXiv:2408.05126
-
LLMJudge: LLMs for Relevance Judgments 9 Aug 2024 · 1 repository · arXiv:2408.08896
-
Can GPT-4 Models Detect Misleading Visualizations? 8 Aug 2024 · 0 repositories · arXiv:2408.12617
-
Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles 8 Aug 2024 · 0 repositories · arXiv:2408.04686
-
Towards Explainable Network Intrusion Detection using Large Language Models 8 Aug 2024 · 0 repositories · arXiv:2408.04342
-
A Comparison of LLM Finetuning Methods & Evaluation Metrics with Travel Chatbot Use Case 7 Aug 2024 · 0 repositories · arXiv:2408.03562
-
Can Rule-Based Insights Enhance LLMs for Radiology Report Classification? Introducing the RadPrompt Methodology 7 Aug 2024 · 0 repositories · arXiv:2408.04121
-
Could ChatGPT get an Engineering Degree? Evaluating Higher Education Vulnerability to AI Assistants 7 Aug 2024 · 0 repositories · arXiv:2408.11841
-
FMiFood: Multi-modal Contrastive Learning for Food Image Classification 7 Aug 2024 · 0 repositories · arXiv:2408.03922
-
Intermediate direct preference optimization 6 Aug 2024 · 0 repositories · arXiv:2408.02923
-
Analysis of Argument Structure Constructions in a Deep Recurrent Language Model 6 Aug 2024 · 0 repositories · arXiv:2408.03062
-
Can LLMs Serve As Time Series Anomaly Detectors? 6 Aug 2024 · 0 repositories · arXiv:2408.03475
-
LLM-Aided Compilation for Tensor Accelerators 6 Aug 2024 · 0 repositories · arXiv:2408.03408
-
LLM-based MOFs Synthesis Condition Extraction using Few-Shot Demonstrations 6 Aug 2024 · 0 repositories · arXiv:2408.04665
-
Is Large Language Model Good at Database Knob Tuning? A Comprehensive Experimental Evaluation 5 Aug 2024 · 0 repositories · arXiv:2408.02213
-
Do Large Language Models Speak All Languages Equally? A Comparative Study in Low-Resource Settings 5 Aug 2024 · 0 repositories · arXiv:2408.02237
-
Why Are My Prompts Leaked? Unraveling Prompt Extraction Threats in Customized Large Language Models 5 Aug 2024 · 1 repository · arXiv:2408.02416
-
SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models 5 Aug 2024 · 1 repository · arXiv:2408.02632
-
Self-Taught Evaluators 5 Aug 2024 · 0 repositories · arXiv:2408.02666
-
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models 4 Aug 2024 · 1 repository · arXiv:2408.01933Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
MedSyn: LLM-based Synthetic Medical Text Generation Framework 4 Aug 2024 · 1 repository · arXiv:2408.02056
-
Advancing Mental Health Pre-Screening: A New Custom GPT for Psychological Distress Assessment 3 Aug 2024 · 0 repositories · arXiv:2408.01614
-
Self-Emotion Blended Dialogue Generation in Social Simulation Agents 3 Aug 2024 · 0 repositories · arXiv:2408.01633
-
Stimulating Imagination: Towards General-purpose Object Rearrangement 3 Aug 2024 · 0 repositories · arXiv:2408.01655
-
A Novel Evaluation Framework for Image2Text Generation 3 Aug 2024 · 0 repositories · arXiv:2408.01723
-
MALADE: Orchestration of LLM-powered Agents with Retrieval Augmented Generation for Pharmacovigilance 3 Aug 2024 · 1 repository · arXiv:2408.01869
-
LLM as Runtime Error Handler: A Promising Pathway to Adaptive Self-Healing of Software Systems 2 Aug 2024 · 0 repositories · arXiv:2408.01055
-
High-Throughput Phenotyping of Clinical Text Using Large Language Models 2 Aug 2024 · 0 repositories · arXiv:2408.01214
-
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation 1 Aug 2024 · 1 repository · arXiv:2408.00764Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Granting GPT-4 License and Opportunity: Enhancing Accuracy and Confidence Estimation for Few-Shot Event Detection 1 Aug 2024 · 0 repositories · arXiv:2408.00914
-
Hybrid Querying Over Relational Databases and Large Language Models 1 Aug 2024 · 0 repositories · arXiv:2408.00884
-
Multi-Level Querying using A Knowledge Pyramid 31 Jul 2024 · 0 repositories · arXiv:2407.21276
-
Interpreting and learning voice commands with a Large Language Model for a robot system 31 Jul 2024 · 0 repositories · arXiv:2407.21512
-
Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation 31 Jul 2024 · 0 repositories · arXiv:2407.21531
-
The Llama 3 Herd of Models 31 Jul 2024 · 5 repositories · arXiv:2407.21783Syntology 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification 30 Jul 2024 · 0 repositories · arXiv:2407.20859
-
Enhancing Agricultural Machinery Management through Advanced LLM Integration 30 Jul 2024 · 0 repositories · arXiv:2407.20588
-
Mimicking the Mavens: Agent-based Opinion Synthesis and Emotion Prediction for Social Media Influencers 30 Jul 2024 · 0 repositories · arXiv:2407.20668
-
SynthVLM: High-Efficiency and High-Quality Synthetic Data for Vision Language Models 30 Jul 2024 · 1 repository · arXiv:2407.20756
-
Improving Retrieval Augmented Language Model with Self-Reasoning 29 Jul 2024 · 0 repositories · arXiv:2407.19813
-
Legal Minds, Algorithmic Decisions: How LLMs Apply Constitutional Principles in Complex Scenarios 29 Jul 2024 · 0 repositories · arXiv:2407.19760
-
Revolutionizing Urban Safety Perception Assessments: Integrating Multimodal Large Language Models with Street View Images 29 Jul 2024 · 0 repositories · arXiv:2407.19719
-
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models 29 Jul 2024 · 0 repositories · arXiv:2407.19914
-
What if Red Can Talk? Dynamic Dialogue Generation Using Large Language Models 29 Jul 2024 · 0 repositories · arXiv:2407.20382
-
Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification on the DAIC-WOZ 27 Jul 2024 · 0 repositories · arXiv:2407.19340
-
GPT Deciphering Fedspeak: Quantifying Dissent Among Hawks and Doves 26 Jul 2024 · 1 repository · arXiv:2407.19110
-
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation 26 Jul 2024 · 1 repository · arXiv:2407.19056Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
TAGIFY: LLM-powered Tagging Interface for Improved Data Findability on OGD portals 26 Jul 2024 · 0 repositories · arXiv:2407.18764
-
Using GPT-4 to guide causal machine learning 26 Jul 2024 · 0 repositories · arXiv:2407.18607
-
Using Large Language Models for the Interpretation of Building Regulations 26 Jul 2024 · 0 repositories · arXiv:2407.21060
-
Cost-effective Instruction Learning for Pathology Vision and Language Analysis 25 Jul 2024 · 1 repository · arXiv:2407.17734
-
Is the Digital Forensics and Incident Response Pipeline Ready for Text-Based Threats in LLM Era? 25 Jul 2024 · 0 repositories · arXiv:2407.17870
-
PersonaGym: Evaluating Persona Agents and LLMs 25 Jul 2024 · 1 repository · arXiv:2407.18416
-
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning 25 Jul 2024 · 1 repository · arXiv:2407.18248Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement 25 Jul 2024 · 0 repositories · arXiv:2407.18370
-
Bailicai: A Domain-Optimized Retrieval-Augmented Generation Framework for Medical Applications 24 Jul 2024 · 0 repositories · arXiv:2407.21055
-
I Could've Asked That: Reformulating Unanswerable Questions 24 Jul 2024 · 1 repository · arXiv:2407.17469
-
Testing Large Language Models on Driving Theory Knowledge and Skills for Connected Autonomous Vehicles 24 Jul 2024 · 0 repositories · arXiv:2407.17211
-
Artificial Intelligence in Extracting Diagnostic Data from Dental Records 23 Jul 2024 · 0 repositories · arXiv:2407.21050
-
Do LLMs Know When to NOT Answer? Investigating Abstention Abilities of Large Language Models 23 Jul 2024 · 0 repositories · arXiv:2407.16221
-
LawLuo: A Multi-Agent Collaborative Framework for Multi-Round Chinese Legal Consultation 23 Jul 2024 · 0 repositories · arXiv:2407.16252
-
Lawma: The Power of Specialization for Legal Tasks 23 Jul 2024 · 0 repositories · arXiv:2407.16615Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection 23 Jul 2024 · 1 repository · arXiv:2407.16237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Patched RTC: evaluating LLMs for diverse software development tasks 23 Jul 2024 · 1 repository · arXiv:2407.16557
-
RedAgent: Red Teaming Large Language Models with Context-aware Autonomous Language Agent 23 Jul 2024 · 0 repositories · arXiv:2407.16667
-
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach 23 Jul 2024 · 0 repositories · arXiv:2407.16833
-
Robust Privacy Amidst Innovation with Large Language Models Through a Critical Assessment of the Risks 23 Jul 2024 · 1 repository · arXiv:2407.16166
-
Can GPT-4 learn to analyse moves in research article abstracts? 22 Jul 2024 · 0 repositories · arXiv:2407.15612
-
Dissecting Multiplication in Transformers: Insights into LLMs 22 Jul 2024 · 1 repository · arXiv:2407.15360
-
Imposter.AI: Adversarial Attacks with Hidden Intentions towards Aligned Large Language Models 22 Jul 2024 · 0 repositories · arXiv:2407.15399
-
MoRSE: Bridging the Gap in Cybersecurity Expertise with Retrieval Augmented Generation 22 Jul 2024 · 0 repositories · arXiv:2407.15748
-
RadioRAG: Factual large language models for enhanced diagnostics in radiology using online retrieval augmented generation 22 Jul 2024 · 1 repository · arXiv:2407.15621
-
Unlocking the Potential: Benchmarking Large Language Models in Water Engineering and Research 22 Jul 2024 · 0 repositories · arXiv:2407.21045
-
Arondight: Red Teaming Large Vision Language Models with Auto-generated Multi-modal Jailbreak Prompts 21 Jul 2024 · 0 repositories · arXiv:2407.15050
-
Toward Adaptive Reasoning in Large Language Models with Thought Rollback 21 Jul 2024 · 1 repository
-
Improving Context-Aware Preference Modeling for Language Models 20 Jul 2024 · 0 repositories · arXiv:2407.14916