Methods › Natural Language Processing › Language Models › GPT-4 › Papers where code ran, page 5
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 5 of 6: papers 401 to 500 of the 526 tagged papers where Syntology ran at least one harvested sample (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4 29 Sep 2023 · 1 repository · arXiv:2309.17277Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving 29 Sep 2023 · 1 repository · arXiv:2309.17452Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 9 pointer-only (licence)
-
LawBench: Benchmarking Legal Knowledge of Large Language Models 28 Sep 2023 · 1 repository · arXiv:2309.16289Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 28 Sep 2023 · 1 repository · arXiv:2309.16583Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs 22 Sep 2023 · 2 repositories · arXiv:2309.13007Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Code Soliloquies for Accurate Calculations in Large Language Models 21 Sep 2023 · 1 repository · arXiv:2309.12161Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" 21 Sep 2023 · 2 repositories · arXiv:2309.12288Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Exploring Iterative Enhancement for Improving Learnersourced Multiple-Choice Question Explanations with Large Language Models 19 Sep 2023 · 1 repository · arXiv:2309.10444Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
InvestLM: A Large Language Model for Investment using Financial Domain Instruction Tuning 15 Sep 2023 · 1 repository · arXiv:2309.13064Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
SafetyBench: Evaluating the Safety of Large Language Models 13 Sep 2023 · 1 repository · arXiv:2309.07045Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
RAIN: Your Language Models Can Align Themselves without Finetuning 13 Sep 2023 · 1 repository · arXiv:2309.07124Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large Language Models for Automated Open-domain Scientific Hypotheses Discovery 6 Sep 2023 · 1 repository · arXiv:2309.02726Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
GPT Can Solve Mathematical Problems Without a Calculator 6 Sep 2023 · 1 repository · arXiv:2309.03241Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes 1 Sep 2023 · 1 repository · arXiv:2309.00237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
BioCoder: A Benchmark for Bioinformatics Code Generation with Large Language Models 31 Aug 2023 · 1 repository · arXiv:2308.16458Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Exploring Large Language Models for Knowledge Graph Completion 26 Aug 2023 · 1 repository · arXiv:2308.13916Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
SciEval: A Multi-Level Large Language Model Evaluation Benchmark for Scientific Research 25 Aug 2023 · 2 repositories · arXiv:2308.13149Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs 25 Aug 2023 · 1 repository · arXiv:2308.13387Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
MLLM-DataEngine: An Iterative Refinement Approach for MLLM 25 Aug 2023 · 1 repository · arXiv:2308.13566Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 1 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
VIGC: Visual Instruction Generation and Correction 24 Aug 2023 · 2 repositories · arXiv:2308.12714Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
InstructionGPT-4: A 200-Instruction Paradigm for Fine-Tuning MiniGPT-4 23 Aug 2023 · 3 repositories · arXiv:2308.12067Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Out of the Cage: How Stochastic Parrots Win in Cyber Security Environments 23 Aug 2023 · 1 repository · arXiv:2308.12086Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Evaluating Large Language Models on Graphs: Performance Insights and Comparative Analysis 22 Aug 2023 · 1 repository · arXiv:2308.11224Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
On the Adversarial Robustness of Multi-Modal Foundation Models 21 Aug 2023 · 1 repository · arXiv:2308.10741Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
LatEval: An Interactive LLMs Evaluation Benchmark with Incomplete Information from Lateral Thinking Puzzles 21 Aug 2023 · 1 repository · arXiv:2308.10855Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
"Guinea Pig Trials" Utilizing GPT: A Novel Smart Agent-Based Modeling Approach for Studying Firm Competition and Collusion 21 Aug 2023 · 2 repositories · arXiv:2308.10974Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ExpeL: LLM Agents Are Experiential Learners 20 Aug 2023 · 2 repositories · arXiv:2308.10144Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data 20 Aug 2023 · 1 repository · arXiv:2308.10253Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Can ChatGPT replace StackOverflow? A Study on Robustness and Reliability of Large Language Model Code Generation 20 Aug 2023 · 1 repository · arXiv:2308.10335Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models 19 Aug 2023 · 1 repository · arXiv:2308.09975Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct 18 Aug 2023 · 1 repository · arXiv:2308.09583Syntology 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models 17 Aug 2023 · 1 repository · arXiv:2308.09729Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Time Travel in LLMs: Tracing Data Contamination in Large Language Models 16 Aug 2023 · 1 repository · arXiv:2308.08493Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Dialogue for Prompting: a Policy-Gradient-Based Discrete Prompt Generation for Few-shot Learning 14 Aug 2023 · 1 repository · arXiv:2308.07272Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher 12 Aug 2023 · 1 repository · arXiv:2308.06463Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
VisIT-Bench: A Benchmark for Vision-Language Instruction Following Inspired by Real-World Use 12 Aug 2023 · 1 repository · arXiv:2308.06595Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Learning Deductive Reasoning from Synthetic Corpus based on Formal Logic 11 Aug 2023 · 3 repositories · arXiv:2308.07336Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs 7 Aug 2023 · 1 repository · arXiv:2308.03349Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Emotionally Numb or Empathetic? Evaluating How LLMs Feel Using EmotionBench 7 Aug 2023 · 1 repository · arXiv:2308.03656Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
"Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models 7 Aug 2023 · 2 repositories · arXiv:2308.03825Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Flows: Building Blocks of Reasoning and Collaborating AI 2 Aug 2023 · 2 repositories · arXiv:2308.01285Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
ChatHome: Development and Evaluation of a Domain-Specific Language Model for Home Renovation 28 Jul 2023 · 1 repository · arXiv:2307.15290Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 10 unverified (of 17 harvested samples) · 1 pointer-only (licence)
-
Enhancing CLIP with GPT-4: Harnessing Visual Descriptions as Prompts 21 Jul 2023 · 1 repository · arXiv:2307.11661Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
L-Eval: Instituting Standardized Evaluation for Long Context Language Models 20 Jul 2023 · 3 repositories · arXiv:2307.11088Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
How is ChatGPT's behavior changing over time? 18 Jul 2023 · 4 repositories · arXiv:2307.09009Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 5 pointer-only (licence)
-
COLLIE: Systematic Construction of Constrained Text Generation Tasks 17 Jul 2023 · 1 repository · arXiv:2307.08689Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
AlpaGasus: Training A Better Alpaca with Fewer Data 17 Jul 2023 · 3 repositories · arXiv:2307.08701Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph 15 Jul 2023 · 3 repositories · arXiv:2307.07697Syntology official (archive's flag): 11 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 15 pointer-only (licence)
-
Leveraging Large Language Models to Generate Answer Set Programs 15 Jul 2023 · 1 repository · arXiv:2307.07699Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Teaching Arithmetic to Small Transformers 7 Jul 2023 · 1 repository · arXiv:2307.03381Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
UMASS_BioNLP at MEDIQA-Chat 2023: Can LLMs generate high-quality synthetic note-oriented doctor-patient conversations? 29 Jun 2023 · 1 repository · arXiv:2306.16931Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding 29 Jun 2023 · 2 repositories · arXiv:2306.17107Syntology official (archive's flag): 1 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 2 pointer-only (licence)
-
LeanDojo: Theorem Proving with Retrieval-Augmented Language Models 27 Jun 2023 · 3 repositories · arXiv:2306.15626Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs 22 Jun 2023 · 1 repository · arXiv:2306.13063Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Joint Prompt Optimization of Stacked LLMs using Variational Inference 21 Jun 2023 · 1 repository · arXiv:2306.12509Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 16 unverified (of 19 harvested samples)
-
ARIES: A Corpus of Scientific Paper Edits Made in Response to Peer Reviews 21 Jun 2023 · 1 repository · arXiv:2306.12587Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena 9 Jun 2023 · 11 repositories · arXiv:2306.05685Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
14 Examples of How LLMs Can Transform Materials Science and Chemistry: A Reflection on a Large Language Model Hackathon 9 Jun 2023 · 2 repositories · arXiv:2306.06283Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases 8 Jun 2023 · 3 repositories · arXiv:2306.05301Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
How Far Can Camels Go? Exploring the State of Instruction Tuning on Open Resources 7 Jun 2023 · 4 repositories · arXiv:2306.04751Syntology official (archive's flag): 5 ran · 6 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
INSTRUCTEVAL: Towards Holistic Evaluation of Instruction-Tuned Large Language Models 7 Jun 2023 · 2 repositories · arXiv:2306.04757Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Benchmarking Large Language Models on CMExam -- A Comprehensive Chinese Medical Exam Dataset 5 Jun 2023 · 1 repository · arXiv:2306.03030Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions 4 Jun 2023 · 1 repository · arXiv:2306.02224Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Self-Verification Improves Few-Shot Clinical Information Extraction 30 May 2023 · 1 repository · arXiv:2306.00024Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
The Rise of AI Language Pathologists: Exploring Two-level Prompt Learning for Few-shot Weakly-supervised Whole Slide Image Classification 29 May 2023 · 1 repository · arXiv:2305.17891Syntology official (archive's flag): 4 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Large Language Models are not Fair Evaluators 29 May 2023 · 1 repository · arXiv:2305.17926Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models 29 May 2023 · 1 repository · arXiv:2305.18189Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Do Language Models Know When They're Hallucinating References? 29 May 2023 · 1 repository · arXiv:2305.18248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated Text 27 May 2023 · 1 repository · arXiv:2305.17359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks 27 May 2023 · 2 repositories · arXiv:2305.17390Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
On Evaluating Adversarial Robustness of Large Vision-Language Models 26 May 2023 · 1 repository · arXiv:2305.16934Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 2 pointer-only (licence)
-
NavGPT: Explicit Reasoning in Vision-and-Language Navigation with Large Language Models 26 May 2023 · 2 repositories · arXiv:2305.16986Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 2 pointer-only (licence)
-
BiomedGPT: A Generalist Vision-Language Foundation Model for Diverse Biomedical Tasks 26 May 2023 · 1 repository · arXiv:2305.17100Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 7 pointer-only (licence)
-
Chain-of-Thought Hub: A Continuous Effort to Measure Large Language Models' Reasoning Performance 26 May 2023 · 1 repository · arXiv:2305.17306Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Self-contradictory Hallucinations of Large Language Models: Evaluation, Detection and Mitigation 25 May 2023 · 1 repository · arXiv:2305.15852Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Voyager: An Open-Ended Embodied Agent with Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16291Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Landmark Attention: Random-Access Infinite Context Length for Transformers 25 May 2023 · 2 repositories · arXiv:2305.16300Syntology official (archive's flag): 1 ran · 11 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
On the Tool Manipulation Capability of Open-source Large Language Models 25 May 2023 · 1 repository · arXiv:2305.16504Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
From Words to Wires: Generating Functioning Electronic Devices from Natural Language Descriptions 24 May 2023 · 1 repository · arXiv:2305.14874Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ByteSized32: A Corpus and Challenge Task for Generating Task-Specific World Models Expressed as Text Games 24 May 2023 · 1 repository · arXiv:2305.14879Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Towards Reliable Misinformation Mitigation: Generalization, Uncertainty, and GPT-4 24 May 2023 · 1 repository · arXiv:2305.14928Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Reasoning with Language Model is Planning with World Model 24 May 2023 · 3 repositories · arXiv:2305.14992Syntology 4 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models 24 May 2023 · 1 repository · arXiv:2305.15064Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models 24 May 2023 · 1 repository · arXiv:2305.15074Syntology official: no sample here; runs from other or unrecorded repositories · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
HuatuoGPT, towards Taming Language Model to Be a Doctor 24 May 2023 · 2 repositories · arXiv:2305.15075Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Peek Across: Improving Multi-Document Modeling via Cross-Document Question-Answering 24 May 2023 · 1 repository · arXiv:2305.15387Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples)
-
Harnessing the Power of Large Language Models for Natural Language to First-Order Logic Translation 24 May 2023 · 1 repository · arXiv:2305.15541Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Aligning Large Language Models through Synthetic Feedback 23 May 2023 · 1 repository · arXiv:2305.13735Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
ZeroSCROLLS: A Zero-Shot Benchmark for Long Text Understanding 23 May 2023 · 1 repository · arXiv:2305.14196Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation 23 May 2023 · 4 repositories · arXiv:2305.14251Syntology official (archive's flag): 10 ran · 11 ran (of which 1 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 6 pointer-only (licence)
-
WikiChat: Stopping the Hallucination of Large Language Model Chatbots by Few-Shot Grounding on Wikipedia 23 May 2023 · 1 repository · arXiv:2305.14292Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
QLoRA: Efficient Finetuning of Quantized LLMs 23 May 2023 · 20 repositories · arXiv:2305.14314Syntology official (archive's flag): 3 ran · 18 ran (of which 1 constructed an object rather than computing a result; 6 with no instrument failure: 2 honoured, 2 violated, 2 with no contract checked; 12 where Syntology's instrument failed) · 8 unverified (of 26 harvested samples) · 17 pointer-only (licence)
-
Dynosaur: A Dynamic Growth Paradigm for Instruction-Tuning Data Curation 23 May 2023 · 1 repository · arXiv:2305.14327Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples)
-
Automatic Model Selection with Large Language Models for Reasoning 23 May 2023 · 1 repository · arXiv:2305.14333Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Large Language Models are Not Yet Human-Level Evaluators for Abstractive Summarization 22 May 2023 · 1 repository · arXiv:2305.13091Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Evaluating the Performance of Large Language Models on GAOKAO Benchmark 21 May 2023 · 1 repository · arXiv:2305.12474Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
TheoremQA: A Theorem-driven Question Answering dataset 21 May 2023 · 1 repository · arXiv:2305.12524Syntology official (archive's flag): 1 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate 19 May 2023 · 1 repository · arXiv:2305.11595Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Generalized Planning in PDDL Domains with Pretrained Large Language Models 18 May 2023 · 1 repository · arXiv:2305.11014Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Tree of Thoughts: Deliberate Problem Solving with Large Language Models 17 May 2023 · 6 repositories · arXiv:2305.10601Syntology official (archive's flag): 2 ran · 19 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 24 harvested samples) · 1 pointer-only (licence)