Methods › Natural Language Processing › Transformers › GPT-3 › Papers where code ran, page 1
GPT-3
Papers archive 2025-07-28
archive papers tagged: 1,906 · with a code link: 866 · where Syntology ran a sample: 319 (259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (319 of 1,906 tagged: 259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 4: papers 1 to 100 of the 319 tagged papers where Syntology ran at least one harvested sample (259 with a run with no instrument failure, 60 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
HealthBench: Evaluating Large Language Models Towards Improved Human Health 13 May 2025 · 1 repository · arXiv:2505.08775Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Context-aware Biases for Length Extrapolation 11 Mar 2025 · 1 repository · arXiv:2503.08067Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
KBQA-o1: Agentic Knowledge Base Question Answering with Monte Carlo Tree Search 31 Jan 2025 · 1 repository · arXiv:2501.18922Syntology official (archive's flag): 2 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 8 harvested samples)
-
Predicting the Performance of Black-box LLMs through Self-Queries 2 Jan 2025 · 1 repository · arXiv:2501.01558Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Training and Evaluating Language Models with Template-based Data Generation 27 Nov 2024 · 1 repository · arXiv:2411.18104Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback 18 Nov 2024 · 1 repository · arXiv:2412.03578Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Fair Summarization: Bridging Quality and Diversity in Extractive Summaries 12 Nov 2024 · 1 repository · arXiv:2411.07521Syntology official (archive's flag): 3 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
SelfCodeAlign: Self-Alignment for Code Generation 31 Oct 2024 · 2 repositories · arXiv:2410.24198Syntology official (archive's flag): 9 ran · 30 ran (of which 3 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 0 violated, 21 with no contract checked; 8 where Syntology's instrument failed) · 7 unverified (of 37 harvested samples)
-
Scaling up Masked Diffusion Models on Text 24 Oct 2024 · 1 repository · arXiv:2410.18514Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 18 Oct 2024 · 1 repository · arXiv:2410.14211Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
The Rise of AI-Generated Content in Wikipedia 10 Oct 2024 · 1 repository · arXiv:2410.08044Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation 5 Oct 2024 · 1 repository · arXiv:2410.04002Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CodeJudge: Evaluating Code Generation with Large Language Models 3 Oct 2024 · 1 repository · arXiv:2410.02184Syntology official (archive's flag): 13 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 2 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 23 harvested samples) · 2 pointer-only (licence)
-
AlignSum: Data Pyramid Hierarchical Fine-tuning for Aligning with Human Summarization Preference 1 Oct 2024 · 1 repository · arXiv:2410.00409Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
MaskLLM: Learnable Semi-Structured Sparsity for Large Language Models 26 Sep 2024 · 1 repository · arXiv:2409.17481Syntology official (archive's flag): 5 ran · 5 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 11 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs 23 Sep 2024 · 1 repository · arXiv:2409.14866Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Neural-Symbolic Collaborative Distillation: Advancing Small Language Models for Complex Reasoning Tasks 20 Sep 2024 · 1 repository · arXiv:2409.13203Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning 18 Sep 2024 · 1 repository · arXiv:2409.12147Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Can Large Language Models Unlock Novel Scientific Research Ideas? 10 Sep 2024 · 1 repository · arXiv:2409.06185Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples)
-
Self-Judge: Selective Instruction Following with Alignment Self-Evaluation 2 Sep 2024 · 1 repository · arXiv:2409.00935Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Unlocking Adversarial Suffix Optimization Without Affirmative Phrases: Efficient Black-box Jailbreaking via LLM as Optimizer 21 Aug 2024 · 1 repository · arXiv:2408.11313Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation 1 Aug 2024 · 1 repository · arXiv:2408.00764Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
PEFT-U: Parameter-Efficient Fine-Tuning for User Personalization 25 Jul 2024 · 1 repository · arXiv:2407.18078Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Enhancing LLM's Cognition via Structurization 23 Jul 2024 · 1 repository · arXiv:2407.16434Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data? 23 Jul 2024 · 1 repository · arXiv:2407.16607Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity 22 Jul 2024 · 1 repository · arXiv:2407.15838Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Does Refusal Training in LLMs Generalize to the Past Tense? 16 Jul 2024 · 1 repository · arXiv:2407.11969Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation 15 Jul 2024 · 1 repository · arXiv:2407.10805Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 1 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples)
-
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct 8 Jul 2024 · 1 repository · arXiv:2407.05700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
SHINE: Saliency-aware HIerarchical NEgative Ranking for Compositional Temporal Grounding 6 Jul 2024 · 1 repository · arXiv:2407.05118Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters 2 Jul 2024 · 1 repository · arXiv:2407.01902Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Language Model Alignment in Multilingual Trolley Problems 2 Jul 2024 · 2 repositories · arXiv:2407.02273Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuning 1 Jul 2024 · 1 repository · arXiv:2407.01320Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 7 pointer-only (licence)
-
ShortcutsBench: A Large-Scale Real-world Benchmark for API-based Agents 28 Jun 2024 · 1 repository · arXiv:2407.00132Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Learning to Plan for Retrieval-Augmented Large Language Models from Knowledge Graphs 20 Jun 2024 · 1 repository · arXiv:2406.14282Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors 20 Jun 2024 · 1 repository · arXiv:2406.14498Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs 20 Jun 2024 · 1 repository · arXiv:2406.14544Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition 19 Jun 2024 · 1 repository · arXiv:2406.13327Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Towards a Client-Centered Assessment of LLM Therapists by Client Simulation 18 Jun 2024 · 1 repository · arXiv:2406.12266Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions 17 Jun 2024 · 1 repository · arXiv:2406.12058Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification 12 Jun 2024 · 1 repository · arXiv:2406.08660Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
BERTs are Generative In-Context Learners 7 Jun 2024 · 1 repository · arXiv:2406.04823Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 13 unverified (of 26 harvested samples)
-
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents 7 Jun 2024 · 2 repositories · arXiv:2406.06613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning 3 Jun 2024 · 1 repository · arXiv:2406.01006Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 15 harvested samples)
-
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction 2 Jun 2024 · 1 repository · arXiv:2406.00755Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Large Language Models are Zero-Shot Next Location Predictors 31 May 2024 · 1 repository · arXiv:2405.20962Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ANAH: Analytical Annotation of Hallucinations in Large Language Models 30 May 2024 · 1 repository · arXiv:2405.20315Syntology official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Aligning to Thousands of Preferences via System Message Generalization 28 May 2024 · 1 repository · arXiv:2405.17977Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
An Empirical Analysis on Large Language Models in Debate Evaluation 28 May 2024 · 1 repository · arXiv:2406.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
THREAD: Thinking Deeper with Recursive Spawning 27 May 2024 · 1 repository · arXiv:2405.17402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
AutoManual: Constructing Instruction Manuals by LLM Agents via Interactive Environmental Learning 25 May 2024 · 1 repository · arXiv:2405.16247Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Evaluating and Safeguarding the Adversarial Robustness of Retrieval-Based In-Context Learning 24 May 2024 · 1 repository · arXiv:2405.15984Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing 23 May 2024 · 1 repository · arXiv:2405.14785Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Eliciting Informative Text Evaluations with Large Language Models 23 May 2024 · 1 repository · arXiv:2405.15077Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment 22 May 2024 · 1 repository · arXiv:2405.13911Syntology official (archive's flag): 3 ran · 6 ran (of which 3 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 3 pointer-only (licence)
-
Zero-Shot Stance Detection using Contextual Data Generation with LLMs 19 May 2024 · 1 repository · arXiv:2405.11637Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis 12 May 2024 · 1 repository · arXiv:2405.07248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Hire Me or Not? Examining Language Model's Behavior with Occupation Attributes 6 May 2024 · 1 repository · arXiv:2405.06687Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Do Large Language Models Understand Conversational Implicature -- A case study with a chinese sitcom 30 Apr 2024 · 1 repository · arXiv:2404.19509Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models 25 Apr 2024 · 1 repository · arXiv:2404.16308Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO 22 Apr 2024 · 1 repository · arXiv:2404.13957Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities 21 Apr 2024 · 1 repository · arXiv:2404.13710Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
From Form(s) to Meaning: Probing the Semantic Depths of Language Models Using Multisense Consistency 18 Apr 2024 · 1 repository · arXiv:2404.12145Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers? 3 Apr 2024 · 1 repository · arXiv:2404.02474Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Advancing LLM Reasoning Generalists with Preference Trees 2 Apr 2024 · 1 repository · arXiv:2404.02078Syntology official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 1 violated, 12 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 20 harvested samples) · 2 pointer-only (licence)
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks 2 Apr 2024 · 1 repository · arXiv:2404.02151Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
FABLES: Evaluating faithfulness and content selection in book-length summarization 1 Apr 2024 · 3 repositories · arXiv:2404.01261Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Vulnerability Detection with Code Language Models: How Far Are We? 27 Mar 2024 · 1 repository · arXiv:2403.18624Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
OmniVid: A Generative Framework for Universal Video Understanding 26 Mar 2024 · 1 repository · arXiv:2403.17935Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
RepairAgent: An Autonomous, LLM-Based Agent for Program Repair 25 Mar 2024 · 1 repository · arXiv:2403.17134Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments 18 Mar 2024 · 1 repository · arXiv:2403.11807Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 2 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models 18 Mar 2024 · 1 repository · arXiv:2403.12171Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Data is all you need: Finetuning LLMs for Chip Design via an Automated design-data augmentation framework 17 Mar 2024 · 1 repository · arXiv:2403.11202Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences 14 Mar 2024 · 2 repositories · arXiv:2403.09032Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought 8 Mar 2024 · 1 repository · arXiv:2403.05518Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Differentially Private Synthetic Data via Foundation Model APIs 2: Text 4 Mar 2024 · 2 repositories · arXiv:2403.01749Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 11 harvested samples)
-
SciAssess: Benchmarking LLM Proficiency in Scientific Literature Analysis 4 Mar 2024 · 1 repository · arXiv:2403.01976Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
SERVAL: Synergy Learning between Vertical Models and LLMs towards Oracle-Level Zero-shot Medical Prediction 3 Mar 2024 · 0 repositories · arXiv:2403.01570Syntology 6 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks 2 Mar 2024 · 1 repository · arXiv:2403.04783Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates 28 Feb 2024 · 1 repository · arXiv:2402.18540Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Measuring Vision-Language STEM Skills of Neural Models 27 Feb 2024 · 1 repository · arXiv:2402.17205Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
ChatMusician: Understanding and Generating Music Intrinsically with LLM 25 Feb 2024 · 1 repository · arXiv:2402.16153Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs 22 Feb 2024 · 1 repository · arXiv:2402.14903Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models 21 Feb 2024 · 1 repository · arXiv:2402.13457Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking Retrieval-Augmented Generation for Medicine 20 Feb 2024 · 2 repositories · arXiv:2402.13178Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
A Critical Evaluation of AI Feedback for Aligning Large Language Models 19 Feb 2024 · 1 repository · arXiv:2402.12366Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement 18 Feb 2024 · 1 repository · arXiv:2402.11436Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
COLD-Attack: Jailbreaking LLMs with Stealthiness and Controllability 13 Feb 2024 · 1 repository · arXiv:2402.08679Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
PRompt Optimization in Multi-Step Tasks (PROMST): Integrating Human Feedback and Heuristic-based Sampling 13 Feb 2024 · 1 repository · arXiv:2402.08702Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs 13 Feb 2024 · 1 repository · arXiv:2402.10962Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Addressing cognitive bias in medical language models 12 Feb 2024 · 1 repository · arXiv:2402.08113Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
CultureLLM: Incorporating Cultural Differences into Large Language Models 9 Feb 2024 · 2 repositories · arXiv:2402.10946Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs 8 Feb 2024 · 2 repositories · arXiv:2402.05668Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Amortized Planning with Large-Scale Transformers: A Case Study on Chess 7 Feb 2024 · 1 repository · arXiv:2402.04494Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Long Is More for Alignment: A Simple but Tough-to-Beat Baseline for Instruction Fine-Tuning 7 Feb 2024 · 1 repository · arXiv:2402.04833Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Training Language Models to Generate Text with Citations via Fine-grained Rewards 6 Feb 2024 · 1 repository · arXiv:2402.04315Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 1 pointer-only (licence)
-
GeReA: Question-Aware Prompt Captions for Knowledge-based Visual Question Answering 4 Feb 2024 · 1 repository · arXiv:2402.02503Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
EffiBench: Benchmarking the Efficiency of Automatically Generated Code 3 Feb 2024 · 1 repository · arXiv:2402.02037Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models 30 Jan 2024 · 1 repository · arXiv:2401.16745Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 2 pointer-only (licence)
-
LLaMP: Large Language Model Made Powerful for High-fidelity Materials Knowledge Retrieval and Distillation 30 Jan 2024 · 1 repository · arXiv:2401.17244Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)