Methods › Natural Language Processing › Language Models › GPT-4 › Papers where code ran, page 2
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 2 of 6: papers 101 to 200 of the 526 tagged papers where Syntology ran at least one harvested sample (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges 18 Jun 2024 · 1 repository · arXiv:2406.12624Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools 18 Jun 2024 · 7 repositories · arXiv:2406.12793Syntology official (archive's flag): 4 ran · 21 ran (of which 0 constructed an object rather than computing a result; 20 with no instrument failure: 0 honoured, 0 violated, 20 with no contract checked; 1 where Syntology's instrument failed) · 8 unverified (of 29 harvested samples) · 1 pointer-only (licence)
-
Can Large Language Models Always Solve Easy Problems if They Can Solve Harder Ones? 18 Jun 2024 · 1 repository · arXiv:2406.12809Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 14 harvested samples) · 4 pointer-only (licence)
-
Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts 18 Jun 2024 · 2 repositories · arXiv:2406.12845Syntology community repositories only · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 18 Jun 2024 · 1 repository · arXiv:2407.13690Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples)
-
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation 17 Jun 2024 · 1 repository · arXiv:2406.11503Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
FinTruthQA: A Benchmark Dataset for Evaluating the Quality of Financial Information Disclosure 17 Jun 2024 · 1 repository · arXiv:2406.12009Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages 14 Jun 2024 · 1 repository · arXiv:2406.09948Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 9 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Neural Concept Binder 14 Jun 2024 · 1 repository · arXiv:2406.09949Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Know the Unknown: An Uncertainty-Sensitive Method for LLM Instruction Tuning 14 Jun 2024 · 1 repository · arXiv:2406.10099Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
JailbreakEval: An Integrated Toolkit for Evaluating Jailbreak Attempts Against Large Language Models 13 Jun 2024 · 1 repository · arXiv:2406.09321Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Judging the Judges: A Systematic Study of Position Bias in LLM-as-a-Judge 12 Jun 2024 · 1 repository · arXiv:2406.07791Syntology official (archive's flag): 16 ran · 16 ran (of which 0 constructed an object rather than computing a result; 16 with no instrument failure: 0 honoured, 0 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 16 harvested samples) · 16 pointer-only (licence)
-
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs 11 Jun 2024 · 1 repository · arXiv:2406.07080Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
AI Sandbagging: Language Models can Strategically Underperform on Evaluations 11 Jun 2024 · 1 repository · arXiv:2406.07358Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
MLLMGuard: A Multi-dimensional Safety Evaluation Suite for Multimodal Large Language Models 11 Jun 2024 · 1 repository · arXiv:2406.07594Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Data-Efficient Learning with Neural Programs 10 Jun 2024 · 1 repository · arXiv:2406.06246Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning 10 Jun 2024 · 1 repository · arXiv:2406.06469Syntology official (archive's flag): 5 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding 8 Jun 2024 · 1 repository · arXiv:2406.05540Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Mixture-of-Agents Enhances Large Language Model Capabilities 7 Jun 2024 · 3 repositories · arXiv:2406.04692Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 3 pointer-only (licence)
-
LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models 7 Jun 2024 · 1 repository · arXiv:2406.05113Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
GameBench: Evaluating Strategic Reasoning Abilities of LLM Agents 7 Jun 2024 · 2 repositories · arXiv:2406.06613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Tool-Planner: Task Planning with Clusters across Multiple Tools 6 Jun 2024 · 1 repository · arXiv:2406.03807Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 6 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
UltraMedical: Building Specialized Generalists in Biomedicine 6 Jun 2024 · 1 repository · arXiv:2406.03949Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Scaling and evaluating sparse autoencoders 6 Jun 2024 · 5 repositories · arXiv:2406.04093Syntology official (archive's flag): 1 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 4 pointer-only (licence)
-
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction 2 Jun 2024 · 1 repository · arXiv:2406.00755Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Phased Instruction Fine-Tuning for Large Language Models 1 Jun 2024 · 1 repository · arXiv:2406.04371Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis 31 May 2024 · 1 repository · arXiv:2405.21075Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 1 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
Query2CAD: Generating CAD models using natural language queries 31 May 2024 · 1 repository · arXiv:2406.00144Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
PATIENT-Ψ: Using Large Language Models to Simulate Patients for Training Mental Health Professionals 30 May 2024 · 1 repository · arXiv:2405.19660Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations 30 May 2024 · 1 repository · arXiv:2405.19740Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples)
-
Preference Alignment with Flow Matching 30 May 2024 · 1 repository · arXiv:2405.19806Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
LLaMEA: A Large Language Model Evolutionary Algorithm for Automatically Generating Metaheuristics 30 May 2024 · 2 repositories · arXiv:2405.20132Syntology official (archive's flag): 3 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning 30 May 2024 · 1 repository · arXiv:2405.20139Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
ANAH: Analytical Annotation of Hallucinations in Large Language Models 30 May 2024 · 1 repository · arXiv:2405.20315Syntology official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
PediatricsGPT: Large Language Models as Chinese Medical Assistants for Pediatric Applications 29 May 2024 · 1 repository · arXiv:2405.19266Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
ORLM: A Customizable Framework in Training Large Models for Automated Optimization Modeling 28 May 2024 · 1 repository · arXiv:2405.17743Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Aligning to Thousands of Preferences via System Message Generalization 28 May 2024 · 1 repository · arXiv:2405.17977Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
An Empirical Analysis on Large Language Models in Debate Evaluation 28 May 2024 · 1 repository · arXiv:2406.00050Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CHESS: Contextual Harnessing for Efficient SQL Synthesis 27 May 2024 · 2 repositories · arXiv:2405.16755Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs 27 May 2024 · 1 repository · arXiv:2405.17013Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
Autoformalizing Euclidean Geometry 27 May 2024 · 1 repository · arXiv:2405.17216Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 2 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples)
-
THREAD: Thinking Deeper with Recursive Spawning 27 May 2024 · 1 repository · arXiv:2405.17402Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Confidence Under the Hood: An Investigation into the Confidence-Probability Alignment in Large Language Models 25 May 2024 · 1 repository · arXiv:2405.16282Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified; the one sample that ran constructed an object rather than computing a result (of 6 harvested samples)
-
STRIDE: A Tool-Assisted LLM Agent Framework for Strategic and Interactive Decision-Making 25 May 2024 · 1 repository · arXiv:2405.16376Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Before Generation, Align it! A Novel and Effective Strategy for Mitigating Hallucinations in Text-to-SQL Generation 24 May 2024 · 1 repository · arXiv:2405.15307Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
JiuZhang3.0: Efficiently Improving Mathematical Reasoning by Training Small Data Synthesis Models 23 May 2024 · 1 repository · arXiv:2405.14365Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
AGILE: A Novel Reinforcement Learning Framework of LLM Agents 23 May 2024 · 1 repository · arXiv:2405.14751Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Can AI Relate: Testing Large Language Model Response for Mental Health Support 20 May 2024 · 1 repository · arXiv:2405.12021Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Observational Scaling Laws and the Predictability of Language Model Performance 17 May 2024 · 1 repository · arXiv:2405.10938Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing 13 May 2024 · 1 repository · arXiv:2405.07726Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
Limited Ability of LLMs to Simulate Human Psychological Behaviours: a Psychometric Analysis 12 May 2024 · 1 repository · arXiv:2405.07248Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MedConceptsQA: Open Source Medical Concepts QA Benchmark 12 May 2024 · 1 repository · arXiv:2405.07348Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Smurfs: Leveraging Multiple Proficiency Agents with Context-Efficiency for Tool Planning 9 May 2024 · 1 repository · arXiv:2405.05955Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli 5 May 2024 · 1 repository · arXiv:2405.02814Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Automating the Enterprise with Foundation Models 3 May 2024 · 1 repository · arXiv:2405.03710Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Self-Play Preference Optimization for Language Model Alignment 1 May 2024 · 1 repository · arXiv:2405.00675Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Do Large Language Models Understand Conversational Implicature -- A case study with a chinese sitcom 30 Apr 2024 · 1 repository · arXiv:2404.19509Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
RepEval: Effective Text Evaluation with LLM Representation 30 Apr 2024 · 1 repository · arXiv:2404.19563Syntology official: no sample here; runs from other or unrecorded repositories · 18 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 2 honoured, 0 violated, 13 with no contract checked; 3 where Syntology's instrument failed) · 9 unverified (of 27 harvested samples) · 6 pointer-only (licence)
-
Constrained Decoding for Secure Code Generation 30 Apr 2024 · 2 repositories · arXiv:2405.00218Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 1 pointer-only (licence)
-
LoRA Land: 310 Fine-tuned LLMs that Rival GPT-4, A Technical Report 29 Apr 2024 · 1 repository · arXiv:2405.00732Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
ComposerX: Multi-Agent Symbolic Music Composition with LLMs 28 Apr 2024 · 1 repository · arXiv:2404.18081Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering 24 Apr 2024 · 1 repository · arXiv:2404.15655Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO 22 Apr 2024 · 1 repository · arXiv:2404.13957Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
SVGEditBench: A Benchmark Dataset for Quantitative Assessment of LLM's SVG Editing Capabilities 21 Apr 2024 · 1 repository · arXiv:2404.13710Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Large Language Models as Test Case Generators: Performance Evaluation and Enhancement 20 Apr 2024 · 1 repository · arXiv:2404.13340Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
CyberSecEval 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models 19 Apr 2024 · 1 repository · arXiv:2404.13161Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective Intelligence 18 Apr 2024 · 1 repository · arXiv:2404.11826Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Uncovering Safety Risks of Large Language Models through Concept Activation Vector 18 Apr 2024 · 1 repository · arXiv:2404.12038Syntology official (archive's flag): 1 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 5 pointer-only (licence)
-
Self-Supervised Visual Preference Alignment 16 Apr 2024 · 1 repository · arXiv:2404.10501Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 1 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents 16 Apr 2024 · 2 repositories · arXiv:2404.10774Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples)
-
Grounded Language Agent for Product Search via Intelligent Web Interactions 16 Apr 2024 · 1 repository · arXiv:2404.10887Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Dataset Reset Policy Optimization for RLHF 12 Apr 2024 · 1 repository · arXiv:2404.08495Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
From Words to Numbers: Your Large Language Model Is Secretly A Capable Regressor When Given In-Context Examples 11 Apr 2024 · 1 repository · arXiv:2404.07544Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation 11 Apr 2024 · 1 repository · arXiv:2404.07917Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models 8 Apr 2024 · 1 repository · arXiv:2404.05221Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Evaluating LLMs at Detecting Errors in LLM Responses 4 Apr 2024 · 1 repository · arXiv:2404.03602Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
AutoWebGLM: A Large Language Model-based Web Navigating Agent 4 Apr 2024 · 1 repository · arXiv:2404.03648Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers? 3 Apr 2024 · 1 repository · arXiv:2404.02474Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models 3 Apr 2024 · 1 repository · arXiv:2404.02823Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks 2 Apr 2024 · 1 repository · arXiv:2404.02151Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-World Code Repositories 31 Mar 2024 · 1 repository · arXiv:2404.00599Syntology official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 14 with no instrument failure: 3 honoured, 0 violated, 11 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 15 harvested samples) · 4 pointer-only (licence)
-
How Much are Large Language Models Contaminated? A Comprehensive Survey and the LLMSanitize Library 31 Mar 2024 · 1 repository · arXiv:2404.00699Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Can LLMs Master Math? Investigating Large Language Models on Math Stack Exchange 30 Mar 2024 · 1 repository · arXiv:2404.00344Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
MANGO: A Benchmark for Evaluating Mapping and Navigation Abilities of Large Language Models 29 Mar 2024 · 1 repository · arXiv:2403.19913Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Enhancing the General Agent Capabilities of Low-Parameter LLMs through Tuning and Multi-Branch Reasoning 29 Mar 2024 · 1 repository · arXiv:2403.19962Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
On-the-fly Definition Augmentation of LLMs for Biomedical NER 29 Mar 2024 · 1 repository · arXiv:2404.00152Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 7 harvested samples)
-
BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text 27 Mar 2024 · 1 repository · arXiv:2403.18421Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Vulnerability Detection with Code Language Models: How Far Are We? 27 Mar 2024 · 1 repository · arXiv:2403.18624Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples)
-
Long-form factuality in large language models 27 Mar 2024 · 3 repositories · arXiv:2403.18802Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models 27 Mar 2024 · 2 repositories · arXiv:2403.18814Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 1 pointer-only (licence)
-
InternLM2 Technical Report 26 Mar 2024 · 3 repositories · arXiv:2403.17297Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
State Space Models as Foundation Models: A Control Theoretic Overview 25 Mar 2024 · 1 repository · arXiv:2403.16899Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 9 harvested samples) · 7 pointer-only (licence)
-
Construction of a Japanese Financial Benchmark for Large Language Models 22 Mar 2024 · 1 repository · arXiv:2403.15062Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations 19 Mar 2024 · 1 repository · arXiv:2403.12451Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning 19 Mar 2024 · 1 repository · arXiv:2403.13164Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments 18 Mar 2024 · 1 repository · arXiv:2403.11807Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 1 honoured, 2 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models 18 Mar 2024 · 1 repository · arXiv:2403.12171Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
CodeUltraFeedback: An LLM-as-a-Judge Dataset for Aligning Large Language Models to Coding Preferences 14 Mar 2024 · 2 repositories · arXiv:2403.09032Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study 13 Mar 2024 · 2 repositories · arXiv:2403.08604Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 3 pointer-only (licence)
-
StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models 12 Mar 2024 · 4 repositories · arXiv:2403.07714Syntology official (archive's flag): 3 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 5 pointer-only (licence)