Methods › Natural Language Processing › Language Models › GPT-4 › Papers where code ran, page 1
GPT-4
Papers archive 2025-07-28
archive papers tagged: 2,870 · with a code link: 1,244 · where Syntology ran a sample: 526 (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (526 of 2,870 tagged: 417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not isolate this method inside it.
Page 1 of 6: papers 1 to 100 of the 526 tagged papers where Syntology ran at least one harvested sample (417 with a run with no instrument failure, 109 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Agent KB: Leveraging Cross-Domain Experience for Agentic Problem Solving 8 Jul 2025 · 1 repository · arXiv:2507.06229Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
DiscoSG: Towards Discourse-Level Text Scene Graph Parsing through Iterative Graph Refinement 18 Jun 2025 · 2 repositories · arXiv:2506.15583Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 1 violated, 16 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 19 harvested samples) · 19 pointer-only (licence)
-
ImpliRet: Benchmarking the Implicit Fact Retrieval Challenge 17 Jun 2025 · 1 repository · arXiv:2506.14407Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
REARANK: Reasoning Re-ranking Agent via Reinforcement Learning 26 May 2025 · 1 repository · arXiv:2505.20046Syntology official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 18 harvested samples) · 4 pointer-only (licence)
-
CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation 20 May 2025 · 1 repository · arXiv:2505.14646Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Lost in Transmission: When and Why LLMs Fail to Reason Globally 13 May 2025 · 0 repositories · arXiv:2505.08140Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
HealthBench: Evaluating Large Language Models Towards Improved Human Health 13 May 2025 · 1 repository · arXiv:2505.08775Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Know Me, Respond to Me: Benchmarking LLMs for Dynamic User Profiling and Personalized Responses at Scale 19 Apr 2025 · 1 repository · arXiv:2504.14225Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
RARE: Retrieval-Augmented Reasoning Modeling 30 Mar 2025 · 1 repository · arXiv:2503.23513Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples)
-
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation 24 Mar 2025 · 1 repository · arXiv:2503.18476Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making 21 Mar 2025 · 0 repositories · arXiv:2503.16965Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
PENCIL: Long Thoughts with Short Memory 18 Mar 2025 · 1 repository · arXiv:2503.14337Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1 13 Mar 2025 · 1 repository · arXiv:2503.10635Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents 12 Mar 2025 · 1 repository · arXiv:2503.09780Syntology official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human? 13 Feb 2025 · 1 repository · arXiv:2502.09416Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Fino1: On the Transferability of Reasoning Enhanced LLMs to Finance 12 Feb 2025 · 1 repository · arXiv:2502.08127Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
SedarEval: Automated Evaluation using Self-Adaptive Rubrics 26 Jan 2025 · 1 repository · arXiv:2501.15595Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Episodic Memories Generation and Evaluation Benchmark for Large Language Models 21 Jan 2025 · 1 repository · arXiv:2501.13121Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 14 pointer-only (licence)
-
PaSa: An LLM Agent for Comprehensive Academic Paper Search 17 Jan 2025 · 1 repository · arXiv:2501.10120Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions 3 Jan 2025 · 1 repository · arXiv:2501.01872Syntology official: no sample here; runs from other or unrecorded repositories · 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 1 pointer-only (licence)
-
Toward Adaptive Reasoning in Large Language Models with Thought Rollback 27 Dec 2024 · 1 repository · arXiv:2412.19707Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Towards Learning to Reason: Comparing LLMs with Neuro-Symbolic on Arithmetic Relations in Abstract Reasoning 7 Dec 2024 · 2 repositories · arXiv:2412.05586Syntology official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Training and Evaluating Language Models with Template-based Data Generation 27 Nov 2024 · 1 repository · arXiv:2411.18104Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation 26 Nov 2024 · 2 repositories · arXiv:2411.17945Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data 22 Nov 2024 · 1 repository · arXiv:2411.15004Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels 21 Nov 2024 · 1 repository · arXiv:2411.13775Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and A Comprehensive Multimodal Dataset Towards General Medical AI 21 Nov 2024 · 1 repository · arXiv:2411.14522Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
PerfCodeGen: Improving Performance of LLM Generated Code with Execution Feedback 18 Nov 2024 · 1 repository · arXiv:2412.03578Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Large Language Models Can Self-Improve in Long-context Reasoning 12 Nov 2024 · 1 repository · arXiv:2411.08147Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
HourVideo: 1-Hour Video-Language Understanding 7 Nov 2024 · 1 repository · arXiv:2411.04998Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
Customized Multiple Clustering via Multi-Modal Subspace Proxy Learning 6 Nov 2024 · 1 repository · arXiv:2411.03978Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Self-Evolved Reward Learning for LLMs 1 Nov 2024 · 1 repository · arXiv:2411.00418Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
RSL-SQL: Robust Schema Linking in Text-to-SQL Generation 31 Oct 2024 · 1 repository · arXiv:2411.00073Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations 30 Oct 2024 · 0 repositories · arXiv:2410.22821Syntology 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
SciPIP: An LLM-based Scientific Paper Idea Proposer 30 Oct 2024 · 1 repository · arXiv:2410.23166Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts 29 Oct 2024 · 1 repository · arXiv:2410.22143Syntology 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Topic-Conversation Relevance (TCR) Dataset and Benchmarks 29 Oct 2024 · 1 repository · arXiv:2411.00038Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
Little Giants: Synthesizing High-Quality Embedding Data at Scale 24 Oct 2024 · 1 repository · arXiv:2410.18634Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
Back to School: Translation Using Grammar Books 20 Oct 2024 · 1 repository · arXiv:2410.15263Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 6 harvested samples)
-
Paths-over-Graph: Knowledge Graph Empowered Large Language Model Reasoning 18 Oct 2024 · 1 repository · arXiv:2410.14211Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples) · 2 pointer-only (licence)
-
TimeSeriesExam: A time series understanding exam 18 Oct 2024 · 1 repository · arXiv:2410.14752Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Looking Inward: Language Models Can Learn About Themselves by Introspection 17 Oct 2024 · 1 repository · arXiv:2410.13787Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards Cross-Cultural Machine Translation with Retrieval-Augmented Generation from Multilingual Knowledge Graphs 17 Oct 2024 · 0 repositories · arXiv:2410.14057Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Evaluating Morphological Compositional Generalization in Large Language Models 16 Oct 2024 · 1 repository · arXiv:2410.12656Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
MSc-SQL: Multi-Sample Critiquing Small Language Models For Text-To-SQL Translation 16 Oct 2024 · 1 repository · arXiv:2410.12916Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
FormalAlign: Automated Alignment Evaluation for Autoformalization 14 Oct 2024 · 1 repository · arXiv:2410.10135Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework 11 Oct 2024 · 1 repository · arXiv:2410.12855Syntology 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Thought2Text: Text Generation from EEG Signal using Large Language Models (LLMs) 10 Oct 2024 · 1 repository · arXiv:2410.07507Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Benchmarking Agentic Workflow Generation 10 Oct 2024 · 1 repository · arXiv:2410.07869Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.08068Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models 10 Oct 2024 · 1 repository · arXiv:2410.12851Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time 9 Oct 2024 · 1 repository · arXiv:2410.06625Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 3 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Take It Easy: Label-Adaptive Self-Rationalization for Fact Verification and Explanation Generation 5 Oct 2024 · 1 repository · arXiv:2410.04002Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
MARPLE: A Benchmark for Long-Horizon Inference 2 Oct 2024 · 1 repository · arXiv:2410.01926Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
CliMB: An AI-enabled Partner for Clinical Predictive Modeling 30 Sep 2024 · 1 repository · arXiv:2410.03736Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
Post-hoc Reward Calibration: A Case Study on Length Bias 25 Sep 2024 · 1 repository · arXiv:2409.17407Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
PAPILLON: Efficient and Stealthy Fuzz Testing-Powered Jailbreaks for LLMs 23 Sep 2024 · 1 repository · arXiv:2409.14866Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models 21 Sep 2024 · 1 repository · arXiv:2409.13989Syntology official (archive's flag): 17 ran · 17 ran (of which 0 constructed an object rather than computing a result; 17 with no instrument failure: 0 honoured, 0 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 18 harvested samples) · 18 pointer-only (licence)
-
ChangeChat: An Interactive Model for Remote Sensing Change Analysis via Multimodal Instruction Tuning 13 Sep 2024 · 1 repository · arXiv:2409.08582Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Can Large Language Models Unlock Novel Scientific Research Ideas? 10 Sep 2024 · 1 repository · arXiv:2409.06185Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 11 harvested samples)
-
Self-Judge: Selective Instruction Following with Alignment Self-Evaluation 2 Sep 2024 · 1 repository · arXiv:2409.00935Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
LongRecipe: Recipe for Efficient Long Context Generalization in Large Language Models 31 Aug 2024 · 1 repository · arXiv:2409.00509Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action 29 Aug 2024 · 1 repository · arXiv:2409.00138Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 13 harvested samples)
-
Interactive Agents: Simulating Counselor-Client Psychological Counseling via Role-Playing LLM-to-LLM Interactions 28 Aug 2024 · 1 repository · arXiv:2408.15787Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
The Mamba in the Llama: Distilling and Accelerating Hybrid Models 27 Aug 2024 · 2 repositories · arXiv:2408.15237Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 5 unverified (of 14 harvested samples) · 5 pointer-only (licence)
-
Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates 23 Aug 2024 · 1 repository · arXiv:2408.13006Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Towards Evaluating and Building Versatile Large Language Models for Medicine 22 Aug 2024 · 1 repository · arXiv:2408.12547Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Revisiting VerilogEval: A Year of Improvements in Large-Language Models for Hardware Code Generation 20 Aug 2024 · 1 repository · arXiv:2408.11053Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Out-of-distribution generalization via composition: a lens through induction heads in Transformers 18 Aug 2024 · 1 repository · arXiv:2408.09503Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
MAG-SQL: Multi-Agent Generative Approach with Soft Schema Linking and Iterative Sub-SQL Refinement for Text-to-SQL 15 Aug 2024 · 1 repository · arXiv:2408.07930Syntology official (archive's flag): 21 ran · 21 ran (of which 0 constructed an object rather than computing a result; 19 with no instrument failure: 1 honoured, 3 violated, 15 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 25 harvested samples) · 4 pointer-only (licence)
-
DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models 4 Aug 2024 · 1 repository · arXiv:2408.01933Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 1 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation 1 Aug 2024 · 1 repository · arXiv:2408.00764Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples)
-
The Llama 3 Herd of Models 31 Jul 2024 · 5 repositories · arXiv:2407.21783Syntology 9 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 9 harvested samples)
-
Self-Training with Direct Preference Optimization Improves Chain-of-Thought Reasoning 25 Jul 2024 · 1 repository · arXiv:2407.18248Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
OriGen:Enhancing RTL Code Generation with Code-to-Code Augmentation and Self-Reflection 23 Jul 2024 · 1 repository · arXiv:2407.16237Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Lawma: The Power of Specialization for Legal Tasks 23 Jul 2024 · 0 repositories · arXiv:2407.16615Syntology 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Does Refusal Training in LLMs Generalize to the Past Tense? 16 Jul 2024 · 1 repository · arXiv:2407.11969Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Sibyl: Simple yet Effective Agent Framework for Complex Real-world Reasoning 15 Jul 2024 · 1 repository · arXiv:2407.10718Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling 13 Jul 2024 · 1 repository · arXiv:2407.09887Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Benchmarking Language Model Creativity: A Case Study on Code Generation 12 Jul 2024 · 1 repository · arXiv:2407.09007Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training 12 Jul 2024 · 2 repositories · arXiv:2407.09121Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
LitSearch: A Retrieval Benchmark for Scientific Literature Search 10 Jul 2024 · 1 repository · arXiv:2407.18940Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples)
-
Automated Peer Reviewing in Paper SEA: Standardization, Evaluation, and Analysis 9 Jul 2024 · 1 repository · arXiv:2407.12857Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples)
-
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct 8 Jul 2024 · 1 repository · arXiv:2407.05700Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models 5 Jul 2024 · 1 repository · arXiv:2407.04693Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples)
-
TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts 3 Jul 2024 · 1 repository · arXiv:2407.03203Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
SeqAR: Jailbreak LLMs with Sequential Auto-Generated Characters 2 Jul 2024 · 1 repository · arXiv:2407.01902Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
The Art of Saying No: Contextual Noncompliance in Language Models 2 Jul 2024 · 1 repository · arXiv:2407.12043Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 1 pointer-only (licence)
-
Deciphering the Factors Influencing the Efficacy of Chain-of-Thought: Probability, Memorization, and Noisy Reasoning 1 Jul 2024 · 1 repository · arXiv:2407.01687Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Sonnet or Not, Bot? Poetry Evaluation for Large Models and Datasets 27 Jun 2024 · 1 repository · arXiv:2406.18906Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
UniGen: A Unified Framework for Textual Dataset Generation Using Large Language Models 27 Jun 2024 · 1 repository · arXiv:2406.18966Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Themis: A Reference-free NLG Evaluation Language Model with Flexibility and Interpretability 26 Jun 2024 · 1 repository · arXiv:2406.18365Syntology official (archive's flag): 5 ran · 5 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs 26 Jun 2024 · 4 repositories · arXiv:2406.18495Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified; the one sample that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 26 Jun 2024 · 1 repository · arXiv:2406.18629Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
Jailbreaking LLMs with Arabic Transliteration and Arabizi 26 Jun 2024 · 1 repository · arXiv:2406.18725Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Ladder: A Model-Agnostic Framework Boosting LLM-based Machine Translation to the Next Level 22 Jun 2024 · 3 repositories · arXiv:2406.15741Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding 20 Jun 2024 · 1 repository · arXiv:2406.14515Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
WONDERBREAD: A Benchmark for Evaluating Multimodal Foundation Models on Business Process Management Tasks 19 Jun 2024 · 1 repository · arXiv:2406.13264Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding 19 Jun 2024 · 1 repository · arXiv:2406.13807Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
An Investigation of Neuron Activation as a Unified Lens to Explain Chain-of-Thought Eliciting Arithmetic Reasoning of LLMs 18 Jun 2024 · 2 repositories · arXiv:2406.12288Syntology official (archive's flag): 11 ran · 34 ran (of which 0 constructed an object rather than computing a result; 32 with no instrument failure: 0 honoured, 3 violated, 29 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 36 harvested samples) · 12 pointer-only (licence)