Methods › General › Learning Rate Schedules › Linear Warmup With Cosine Annealing › Papers where code ran, page 3
Linear Warmup With Cosine Annealing
Papers archive 2025-07-28
archive papers tagged: 3,797 · with a code link: 1,655 · where Syntology ran a sample: 602 (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (602 of 3,797 tagged: 490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 3 of 7: papers 201 to 300 of the 602 tagged papers where Syntology ran at least one harvested sample (490 with a run with no instrument failure, 112 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Enhancing In-context Learning via Linear Probe Calibration 22 Jan 2024 · 1 repository · arXiv:2401.12406Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
BadChain: Backdoor Chain-of-Thought Prompting for Large Language Models 20 Jan 2024 · 1 repository · arXiv:2401.12242Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs 18 Jan 2024 · 1 repository · arXiv:2401.10065Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning 16 Jan 2024 · 1 repository · arXiv:2401.08326Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 1 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples)
-
Tuning Language Models by Proxy 16 Jan 2024 · 2 repositories · arXiv:2401.08565Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs 12 Jan 2024 · 2 repositories · arXiv:2401.06373Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Mission: Impossible Language Models 12 Jan 2024 · 1 repository · arXiv:2401.06416Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
The Benefits of a Concise Chain of Thought on Problem-Solving in Large Language Models 11 Jan 2024 · 1 repository · arXiv:2401.05618Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning 10 Jan 2024 · 1 repository · arXiv:2401.05268Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples)
-
InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks 10 Jan 2024 · 1 repository · arXiv:2401.05507Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Mixtral of Experts 8 Jan 2024 · 6 repositories · arXiv:2401.04088Syntology 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 5 samples that ran constructed an object rather than computing a result (of 5 harvested samples)
-
LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models 1 Jan 2024 · 1 repository · arXiv:2401.00757Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Jatmo: Prompt Injection Defense by Task-Specific Finetuning 29 Dec 2023 · 1 repository · arXiv:2312.17673Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
SecQA: A Concise Question-Answering Dataset for Evaluating Large Language Models in Computer Security 26 Dec 2023 · 1 repository · arXiv:2312.15838Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Principled Instructions Are All You Need for Questioning LLaMA-1/2, GPT-3.5/4 26 Dec 2023 · 2 repositories · arXiv:2312.16171Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
De novo Drug Design using Reinforcement Learning with Multiple GPT Agents 21 Dec 2023 · 2 repositories · arXiv:2401.06155Syntology official (archive's flag): 3 ran · 3 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation 20 Dec 2023 · 1 repository · arXiv:2312.13010Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
An In-depth Look at Gemini's Language Abilities 18 Dec 2023 · 1 repository · arXiv:2312.11444Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Causality Analysis for Evaluating the Security of Large Language Models 13 Dec 2023 · 1 repository · arXiv:2312.07876Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
AI Control: Improving Safety Despite Intentional Subversion 12 Dec 2023 · 1 repository · arXiv:2312.06942Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Image Content Generation with Causal Reasoning 12 Dec 2023 · 1 repository · arXiv:2312.07132Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Can It Edit? Evaluating the Ability of Large Language Models to Follow Code Editing Instructions 11 Dec 2023 · 1 repository · arXiv:2312.12450Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically 4 Dec 2023 · 2 repositories · arXiv:2312.02119Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Generative Parameter-Efficient Fine-Tuning 1 Dec 2023 · 1 repository · arXiv:2312.00700Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 16 harvested samples) · 10 pointer-only (licence)
-
Robust Concept Erasure via Kernelized Rate-Distortion Maximization 30 Nov 2023 · 1 repository · arXiv:2312.00194Syntology official (archive's flag): 20 ran · 20 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 21 harvested samples) · 2 pointer-only (licence)
-
CharacterGLM: Customizing Chinese Conversational AI Characters with Large Language Models 28 Nov 2023 · 1 repository · arXiv:2311.16832Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation 26 Nov 2023 · 1 repository · arXiv:2311.15296Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Hardware Resilience Properties of Text-Guided Image Classifiers 23 Nov 2023 · 1 repository · arXiv:2311.14062Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
PG-Video-LLaVA: Pixel Grounding Large Video-Language Models 22 Nov 2023 · 1 repository · arXiv:2311.13435Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Evil Geniuses: Delving into the Safety of LLM-based Agents 20 Nov 2023 · 1 repository · arXiv:2311.11855Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Event Causality Is Key to Computational Story Understanding 16 Nov 2023 · 1 repository · arXiv:2311.09648Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
FinanceMath: Knowledge-Intensive Math Reasoning in Finance Domains 16 Nov 2023 · 1 repository · arXiv:2311.09797Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
ToolTalk: Evaluating Tool-Usage in a Conversational Setting 15 Nov 2023 · 1 repository · arXiv:2311.10775Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models 10 Nov 2023 · 2 repositories · arXiv:2311.06233Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Smart Agent-Based Modeling: On the Use of Large Language Models in Computer Simulations 10 Nov 2023 · 4 repositories · arXiv:2311.06330Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Agent Lumos: Unified and Modular Training for Open-Source Language Agents 9 Nov 2023 · 2 repositories · arXiv:2311.05657Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Massive Editing for Large Language Models via Meta Learning 8 Nov 2023 · 1 repository · arXiv:2311.04661Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Rethinking Benchmark and Contamination for Language Models with Rephrased Samples 8 Nov 2023 · 1 repository · arXiv:2311.04850Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Neuro-GPT: Towards A Foundation Model for EEG 7 Nov 2023 · 1 repository · arXiv:2311.03764Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models 7 Nov 2023 · 1 repository · arXiv:2311.04131Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
DeepInception: Hypnotize Large Language Model to Be Jailbreaker 6 Nov 2023 · 1 repository · arXiv:2311.03191Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization 30 Oct 2023 · 1 repository · arXiv:2310.20033Syntology official (archive's flag): 7 ran · 7 ran (of which 3 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
Large language models for aspect-based sentiment analysis 27 Oct 2023 · 1 repository · arXiv:2310.18025Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
LightLM: A Lightweight Deep and Narrow Language Model for Generative Recommendation 26 Oct 2023 · 1 repository · arXiv:2310.17488Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 1 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
In-Context Learning Dynamics with Random Binary Sequences 26 Oct 2023 · 1 repository · arXiv:2310.17639Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution 25 Oct 2023 · 4 repositories · arXiv:2310.16834Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 0 violated, 13 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 18 harvested samples) · 15 pointer-only (licence)
-
TRAMS: Training-free Memory Selection for Long-range Language Modeling 24 Oct 2023 · 1 repository · arXiv:2310.15494Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Learning From Free-Text Human Feedback -- Collect New Datasets Or Extend Existing Ones? 24 Oct 2023 · 1 repository · arXiv:2310.15758Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Towards a Mechanistic Interpretation of Multi-Step Reasoning Capabilities of Language Models 23 Oct 2023 · 2 repositories · arXiv:2310.14491Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Evaluating Spatial Understanding of Large Language Models 23 Oct 2023 · 1 repository · arXiv:2310.14540Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples)
-
Language Models Hallucinate, but May Excel at Fact Verification 23 Oct 2023 · 1 repository · arXiv:2310.14564Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
LLM-in-the-loop: Leveraging Large Language Model for Thematic Analysis 23 Oct 2023 · 1 repository · arXiv:2310.15100Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic Provers 23 Oct 2023 · 1 repository · arXiv:2310.15164Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4 21 Oct 2023 · 1 repository · arXiv:2310.13988Syntology 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Cache me if you Can: an Online Cost-aware Teacher-Student framework to Reduce the Calls to Large Language Models 20 Oct 2023 · 1 repository · arXiv:2310.13395Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A Simple Baseline for Knowledge-Based Visual Question Answering 20 Oct 2023 · 0 repositories · arXiv:2310.13570Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT Interactions 19 Oct 2023 · 1 repository · arXiv:2310.12418Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Identifying and Adapting Transformer-Components Responsible for Gender Bias in an English Language Model 19 Oct 2023 · 1 repository · arXiv:2310.12611Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
AgentTuning: Enabling Generalized Agent Abilities for LLMs 19 Oct 2023 · 1 repository · arXiv:2310.12823Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Probing the Creativity of Large Language Models: Can models produce divergent semantic association? 17 Oct 2023 · 1 repository · arXiv:2310.11158Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Data Contamination Through the Lens of Time 16 Oct 2023 · 1 repository · arXiv:2310.10628Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology 16 Oct 2023 · 1 repository · arXiv:2310.10632Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Approximating Two-Layer Feedforward Networks for Efficient Transformers 16 Oct 2023 · 2 repositories · arXiv:2310.10837Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
QUIK: Towards End-to-End 4-Bit Inference on Generative Large Language Models 13 Oct 2023 · 1 repository · arXiv:2310.09259Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning 13 Oct 2023 · 1 repository · arXiv:2310.09430Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
Jailbreaking Black Box Large Language Models in Twenty Queries 12 Oct 2023 · 1 repository · arXiv:2310.08419Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 5 unverified (of 7 harvested samples)
-
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog 11 Oct 2023 · 2 repositories · arXiv:2310.07259Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Found in the Middle: Permutation Self-Consistency Improves Listwise Ranking in Large Language Models 11 Oct 2023 · 1 repository · arXiv:2310.07712Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Large Language Models Are Zero-Shot Time Series Forecasters 11 Oct 2023 · 2 repositories · arXiv:2310.07820Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 1 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
GeoLLM: Extracting Geospatial Knowledge from Large Language Models 10 Oct 2023 · 1 repository · arXiv:2310.06213Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Humans and language models diverge when predicting repeating text 10 Oct 2023 · 1 repository · arXiv:2310.06408Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT 7 Oct 2023 · 2 repositories · arXiv:2310.04673Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Large Language Models Only Pass Primary School Exams in Indonesia: A Comprehensive Test on IndoMMLU 7 Oct 2023 · 1 repository · arXiv:2310.04928Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models 6 Oct 2023 · 2 repositories · arXiv:2310.04406Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 2 pointer-only (licence)
-
Copy Suppression: Comprehensively Understanding an Attention Head 6 Oct 2023 · 1 repository · arXiv:2310.04625Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
Agent Instructs Large Language Models to be General Zero-Shot Reasoners 5 Oct 2023 · 1 repository · arXiv:2310.03710Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples)
-
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines 5 Oct 2023 · 3 repositories · arXiv:2310.03714Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples)
-
Automating Human Tutor-Style Programming Feedback: Leveraging GPT-4 Tutor Model for Hint Generation and GPT-3.5 Student Model for Hint Validation 5 Oct 2023 · 2 repositories · arXiv:2310.03780Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
NOLA: Compressing LoRA using Linear Combination of Random Basis 4 Oct 2023 · 1 repository · arXiv:2310.02556Syntology official (archive's flag): 5 ran · 5 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture 4 Oct 2023 · 1 repository · arXiv:2310.03052Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 10 harvested samples)
-
Instances Need More Care: Rewriting Prompts for Instances with LLMs in the Loop Yields Better Zero-Shot Performance 3 Oct 2023 · 1 repository · arXiv:2310.02107Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
GPT-Driver: Learning to Drive with GPT 2 Oct 2023 · 1 repository · arXiv:2310.01415Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
LLM Lies: Hallucinations are not Bugs, but Features as Adversarial Examples 2 Oct 2023 · 1 repository · arXiv:2310.01469Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Analyzing and Mitigating Object Hallucination in Large Vision-Language Models 1 Oct 2023 · 1 repository · arXiv:2310.00754Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
BooookScore: A systematic exploration of book-length summarization in the era of LLMs 1 Oct 2023 · 2 repositories · arXiv:2310.00785Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Memory Gym: Towards Endless Tasks to Benchmark Memory Capabilities of Agents 29 Sep 2023 · 1 repository · arXiv:2309.17207Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Cooperation, Competition, and Maliciousness: LLM-Stakeholders Interactive Negotiation 29 Sep 2023 · 2 repositories · arXiv:2309.17234Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
GPT-Fathom: Benchmarking Large Language Models to Decipher the Evolutionary Path towards GPT-4 and Beyond 28 Sep 2023 · 1 repository · arXiv:2309.16583Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
MindGPT: Interpreting What You See with Non-invasive Brain Recordings 27 Sep 2023 · 1 repository · arXiv:2309.15729Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
How to Catch an AI Liar: Lie Detection in Black-Box LLMs by Asking Unrelated Questions 26 Sep 2023 · 1 repository · arXiv:2309.15840Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models 21 Sep 2023 · 1 repository · arXiv:2309.12284Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 7 unverified (of 22 harvested samples)
-
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A" 21 Sep 2023 · 2 repositories · arXiv:2309.12288Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models 20 Sep 2023 · 1 repository · arXiv:2309.11674Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
RECAP: Retrieval-Augmented Audio Captioning 18 Sep 2023 · 1 repository · arXiv:2309.09836Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers 15 Sep 2023 · 2 repositories · arXiv:2309.08532Syntology 11 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 0 violated, 1 with no contract checked; 8 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples)
-
ChatGPT MT: Competitive for High- (but not Low-) Resource Languages 14 Sep 2023 · 2 repositories · arXiv:2309.07423Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Circuit Breaking: Removing Model Behaviors with Targeted Ablation 12 Sep 2023 · 1 repository · arXiv:2309.05973Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models 11 Sep 2023 · 1 repository · arXiv:2309.05605Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Publicly Shareable Clinical Large Language Model Built on Synthetic Clinical Notes 1 Sep 2023 · 1 repository · arXiv:2309.00237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Taken out of context: On measuring situational awareness in LLMs 1 Sep 2023 · 1 repository · arXiv:2309.00667Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)