Methods › Natural Language Processing › Transformers › GPT-3 › Papers, page 8
GPT-3
Papers archive 2025-07-28
archive papers tagged: 1,906 · with a code link: 866 · where Syntology ran a sample: 319 (259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (319 of 1,906 tagged: 259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument)
Page 8 of 20: papers 701 to 800 of 1,906, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Critical Evaluation of AI Feedback for Aligning Large Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12366Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Ask Optimal Questions: Aligning Large Language Models with Retriever's Preference in Conversational Search 19 Feb 2024 · 0 repositories · arXiv:2402.11827
-
DeepCode AI Fix: Fixing Security Vulnerabilities with Large Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.13291
-
Surprising Efficacy of Fine-Tuned Transformers for Fact-Checking over Larger Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.12147
-
Is Open-Source There Yet? A Comparative Study on Commercial and Open-Source LLMs in Their Ability to Label Chest X-Ray Reports 19 Feb 2024 · 0 repositories · arXiv:2402.12298
-
Enabling Weak LLMs to Judge Response Reliability via Meta Ranking 19 Feb 2024 · 0 repositories · arXiv:2402.12146
-
Query-Based Adversarial Prompt Generation 19 Feb 2024 · 2 repositories · arXiv:2402.12329
-
SPML: A DSL for Defending Language Models Against Prompt Attacks 19 Feb 2024 · 0 repositories · arXiv:2402.11755
-
Stick to your Role! Stability of Personal Values Expressed in Large Language Models 19 Feb 2024 · 0 repositories · arXiv:2402.14846
-
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One 19 Feb 2024 · 0 repositories · arXiv:2402.12150
-
Decoding News Narratives: A Critical Analysis of Large Language Models in Framing Detection 18 Feb 2024 · 0 repositories · arXiv:2402.11621
-
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement 18 Feb 2024 · 1 repository · arXiv:2402.11436Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Exploring ChatGPT for Next-generation Information Retrieval: Opportunities and Challenges 17 Feb 2024 · 0 repositories · arXiv:2402.11203
-
GenDec: A robust generative Question-decomposition method for Multi-hop reasoning 17 Feb 2024 · 0 repositories · arXiv:2402.11166
-
Assessing the Reasoning Abilities of ChatGPT in the Context of Claim Verification 16 Feb 2024 · 0 repositories · arXiv:2402.10735
-
Can Separators Improve Chain-of-Thought Prompting? 16 Feb 2024 · 0 repositories · arXiv:2402.10645
-
Disordered-DABS: A Benchmark for Dynamic Aspect-Based Summarization in Disordered Texts 16 Feb 2024 · 1 repository · arXiv:2402.10554
-
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling 16 Feb 2024 · 1 repository · arXiv:2402.10466Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Large Language Models Fall Short: Understanding Complex Relationships in Detective Narratives 16 Feb 2024 · 0 repositories · arXiv:2402.11051
-
LLMs in the Heart of Differential Testing: A Case Study on a Medical Rule Engine 16 Feb 2024 · 0 repositories · arXiv:2404.03664
-
Cultural Commonsense Knowledge for Intercultural Dialogues 16 Feb 2024 · 0 repositories · arXiv:2402.10689
-
Universal Prompt Optimizer for Safe Text-to-Image Generation 16 Feb 2024 · 1 repository · arXiv:2402.10882
-
An Analysis of Language Frequency and Error Correction for Esperanto 15 Feb 2024 · 0 repositories · arXiv:2402.09696
-
Fine-tuning Large Language Model (LLM) Artificial Intelligence Chatbots in Ophthalmology and LLM-based evaluation using GPT-4 15 Feb 2024 · 0 repositories · arXiv:2402.10083
-
PAL: Proxy-Guided Black-Box Attack on Large Language Models 15 Feb 2024 · 1 repository · arXiv:2402.09674
-
The Butterfly Effect of Model Editing: Few Edits Can Trigger Large Language Models Collapse 15 Feb 2024 · 1 repository · arXiv:2402.09656
-
API Pack: A Massive Multi-Programming Language Dataset for API Call Generation 14 Feb 2024 · 1 repository · arXiv:2402.09615
-
MPIrigen: MPI Code Generation through Domain-Specific Language Models 14 Feb 2024 · 1 repository · arXiv:2402.09126
-
"Reasoning" with Rhetoric: On the Style-Evidence Tradeoff in LLM-Generated Counter-Arguments 13 Feb 2024 · 0 repositories · arXiv:2402.08498
-
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability 13 Feb 2024 · 1 repository · arXiv:2402.08679Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Lying Blindly: Bypassing ChatGPT's Safeguards to Generate Hard-to-Detect Disinformation Claims 13 Feb 2024 · 0 repositories · arXiv:2402.08467
-
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs 13 Feb 2024 · 1 repository · arXiv:2402.10962Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Mitigating Object Hallucination in Large Vision-Language Models via Classifier-Free Guidance 13 Feb 2024 · 0 repositories · arXiv:2402.08680
-
PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling 13 Feb 2024 · 1 repository · arXiv:2402.08702Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Addressing cognitive bias in medical language models 12 Feb 2024 · 1 repository · arXiv:2402.08113Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
BreakGPT: A Large Language Model with Multi-stage Structure for Financial Breakout Detection 12 Feb 2024 · 1 repository · arXiv:2402.07536
-
CyberMetric: A Benchmark Dataset based on Retrieval-Augmented Generation for Evaluating LLMs in Cybersecurity Knowledge 12 Feb 2024 · 1 repository · arXiv:2402.07688
-
Investigating the Impact of Data Contamination of Large Language Models in Text-to-SQL Translation 12 Feb 2024 · 0 repositories · arXiv:2402.08100
-
Can Graph Descriptive Order Affect Solving Graph Problems with LLMs? 11 Feb 2024 · 0 repositories · arXiv:2402.07140
-
ChemLLM: A Chemical Large Language Model 10 Feb 2024 · 1 repository · arXiv:2402.06852
-
CultureLLM: Incorporating Cultural Differences into Large Language Models 9 Feb 2024 · 2 repositories · arXiv:2402.10946Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ExaRanker-Open: Synthetic Explanation for IR using Open-Source LLMs 9 Feb 2024 · 1 repository · arXiv:2402.06334
-
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs 8 Feb 2024 · 2 repositories · arXiv:2402.05668Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
In-Context Principle Learning from Mistakes 8 Feb 2024 · 1 repository · arXiv:2402.05403
-
Zero-Shot Chain-of-Thought Reasoning Guided by Evolutionary Algorithms in Large Language Models 8 Feb 2024 · 0 repositories · arXiv:2402.05376
-
A Hypothesis-Driven Framework for the Analysis of Self-Rationalising Models 7 Feb 2024 · 1 repository · arXiv:2402.04787
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 · 1 repository · arXiv:2402.04494Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach 7 Feb 2024 · 0 repositories · arXiv:2402.04609
-
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning 7 Feb 2024 · 1 repository · arXiv:2402.04833Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Behind the Screen: Investigating ChatGPT's Dark Personality Traits and Conspiracy Beliefs 6 Feb 2024 · 0 repositories · arXiv:2402.04110
-
Detecting Mode Collapse in Language Models via Narration 6 Feb 2024 · 0 repositories · arXiv:2402.04477
-
Large Language Models as an Indirect Reasoner: Contrapositive and Contradiction for Automated Reasoning 6 Feb 2024 · 0 repositories · arXiv:2402.03667
-
Large Language Models As MOOCs Graders 6 Feb 2024 · 0 repositories · arXiv:2402.03776
-
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLMs 6 Feb 2024 · 0 repositories · arXiv:2402.03927
-
Are Machines Better at Complex Reasoning? Unveiling Human-Machine Inference Gaps in Entailment Verification 6 Feb 2024 · 0 repositories · arXiv:2402.03686
-
Training Language Models to Generate Text with Citations via Fine-grained Rewards 6 Feb 2024 · 1 repository · arXiv:2402.04315Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT Models 5 Feb 2024 · 1 repository · arXiv:2402.02987
-
Harnessing PubMed User Query Logs for Post Hoc Explanations of Recommended Similar Articles 5 Feb 2024 · 0 repositories · arXiv:2402.03484
-
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models 5 Feb 2024 · 1 repository · arXiv:2402.02896
-
SWAG: Storytelling With Action Guidance 5 Feb 2024 · 1 repository · arXiv:2402.03483
-
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering 4 Feb 2024 · 1 repository · arXiv:2402.02503Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation 4 Feb 2024 · 0 repositories · arXiv:2402.14594
-
EffiBench: Benchmarking the Efficiency of Automatically Generated Code 3 Feb 2024 · 1 repository · arXiv:2402.02037Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Can LLMs perform structured graph reasoning? 2 Feb 2024 · 1 repository · arXiv:2402.01805
-
Learning Planning-based Reasoning by Trajectories Collection and Process Reward Synthesizing 1 Feb 2024 · 0 repositories · arXiv:2402.00658
-
Self-Supervised Contrastive Pre-Training for Multivariate Point Processes 1 Feb 2024 · 0 repositories · arXiv:2402.00987
-
Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization? 1 Feb 2024 · 0 repositories · arXiv:2402.00841
-
Global-Liar: Factuality of LLMs over Time and Geographic Regions 31 Jan 2024 · 0 repositories · arXiv:2401.17839
-
Making a Long Story Short in Conversation Modeling 31 Jan 2024 · 0 repositories · arXiv:2402.00143
-
Mitigating the Influence of Distractor Tasks in LMs with Prior-Aware Decoding 31 Jan 2024 · 0 repositories · arXiv:2401.17692
-
Paramanu: A Family of Novel Efficient Generative Foundation Language Models for Indian Languages 31 Jan 2024 · 0 repositories · arXiv:2401.18034
-
A Preliminary Study on Using Large Language Models in Software Pentesting 30 Jan 2024 · 0 repositories · arXiv:2401.17459
-
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation 30 Jan 2024 · 1 repository · arXiv:2401.17244Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.16745Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
Leveraging Professional Radiologists' Expertise to Enhance LLMs' Evaluation for Radiology Reports 29 Jan 2024 · 0 repositories · arXiv:2401.16578
-
LLM4Vuln: A Unified Evaluation Framework for Decoupling and Enhancing LLMs' Vulnerability Reasoning 29 Jan 2024 · 0 repositories · arXiv:2401.16185
-
ReGAL: Refactoring Programs to Discover Generalizable Abstractions 29 Jan 2024 · 1 repository · arXiv:2401.16467Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
An Insight into Security Code Review with LLMs: Capabilities, Obstacles, and Influential Factors 29 Jan 2024 · 0 repositories · arXiv:2401.16310
-
Enhancing Large Language Model Performance To Answer Questions and Extract Information More Accurately 27 Jan 2024 · 0 repositories · arXiv:2402.01722
-
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance 27 Jan 2024 · 0 repositories · arXiv:2401.15328
-
Fortifying Ethical Boundaries in AI: Advanced Strategies for Enhancing Security in Large Language Models 27 Jan 2024 · 0 repositories · arXiv:2402.01725
-
Scalable Qualitative Coding with LLMs: Chain-of-Thought Reasoning Matches Human Performance in Some Hermeneutic Tasks 26 Jan 2024 · 0 repositories · arXiv:2401.15170
-
A comparative study of zero-shot inference with large language models and supervised modeling in breast cancer pathology classification 25 Jan 2024 · 0 repositories · arXiv:2401.13887
-
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 25 Jan 2024 · 1 repository · arXiv:2401.14196Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Evaluating GPT-3.5's Awareness and Summarization Abilities for European Constitutional Texts with Shared Topics 25 Jan 2024 · 0 repositories · arXiv:2401.14524
-
Investigate-Consolidate-Exploit: A General Strategy for Inter-Task Agent Self-Evolution 25 Jan 2024 · 0 repositories · arXiv:2401.13996
-
LongHealth: A Question Answering Benchmark with Long Clinical Documents 25 Jan 2024 · 1 repository · arXiv:2401.14490Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
TrICy: Trigger-guided Data-to-text Generation with Intent aware Attention-Copy 25 Jan 2024 · 0 repositories · arXiv:2402.01714
-
Unmasking and Quantifying Racial Bias of Large Language Models in Medical Report Generation 25 Jan 2024 · 0 repositories · arXiv:2401.13867
-
ZS4C: Zero-Shot Synthesis of Compilable Code for Incomplete Code Snippets using LLMs 25 Jan 2024 · 0 repositories · arXiv:2401.14279
-
Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4 24 Jan 2024 · 0 repositories · arXiv:2401.13810
-
Can GPT-3.5 Generate and Code Discharge Summaries? 24 Jan 2024 · 1 repository · arXiv:2401.13512Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Evaluation of General Large Language Models in Contextually Assessing Semantic Concepts Extracted from Adult Critical Care Electronic Health Record Notes 24 Jan 2024 · 0 repositories · arXiv:2401.13588
-
KAM-CoT: Knowledge Augmented Multimodal Chain-of-Thoughts Reasoning 23 Jan 2024 · 0 repositories · arXiv:2401.12863
-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models 20 Jan 2024 · 1 repository · arXiv:2401.12242Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Enhancing Large Language Models for Clinical Decision Support by Incorporating Clinical Practice Guidelines 20 Jan 2024 · 0 repositories · arXiv:2401.11120
-
Evaluating and Enhancing Large Language Models Performance in Domain-specific Medicine: Osteoarthritis Management with DocOA 20 Jan 2024 · 0 repositories · arXiv:2401.12998
-
FinLLMs: A Framework for Financial Reasoning Dataset Generation with Large Language Models 19 Jan 2024 · 0 repositories · arXiv:2401.10744
-
Mining experimental data from Materials Science literature with Large Language Models: an evaluation study 19 Jan 2024 · 1 repository · arXiv:2401.11052
-
Gender Bias in Machine Translation and The Era of Large Language Models 18 Jan 2024 · 0 repositories · arXiv:2401.10016