Browse State-of-the-Art › Instruction Following › Papers, page 7
Instruction Following
Papers archive 2025-07-28
archive papers tagged: 1,135 · with a code link: 609 · where Syntology ran a sample: 311 (255 with a run with no instrument failure, 56 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (311 of 1,135 tagged: 255 with a run with no instrument failure, 56 where every run was a failure of Syntology's instrument)
Page 7 of 12: papers 601 to 700 of 1,135, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
15 Apr 2020 1 repository listed
-
1 Dec 2019 1 repository listed
-
21 Oct 2019 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
23 Jul 2019 1 repository listed
-
3 Jul 2019 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
19 Nov 2018 1 repository listed
-
10 Nov 2018 1 repository listed
-
31 May 2018 1 repository listed
-
26 Aug 2015 1 repository listed
-
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning17 Jul 2025 0 repositories listed
-
How Many Instructions Can LLMs Follow at Once?15 Jul 2025 0 repositories listed
-
Multilingual Multimodal Software Developer for Code Generation11 Jul 2025 0 repositories listed
-
TuneShield: Mitigating Toxicity in Conversational AI while Fine-tuning on Untrusted Data8 Jul 2025 0 repositories listed
-
Bridging Offline and Online Reinforcement Learning for LLMs26 Jun 2025 0 repositories listed
-
Multi-lingual Functional Evaluation for Large Language Models25 Jun 2025 0 repositories listed
-
Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models24 Jun 2025 0 repositories listed
-
JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent21 Jun 2025 0 repositories listed
-
Treasure Hunt: Real-time Targeting of the Long Tail using Training-Time Markers17 Jun 2025 0 repositories listed
-
Instruction Following by Boosting Attention of Large Language Models16 Jun 2025 0 repositories listed
-
LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction16 Jun 2025 0 repositories listed
-
Mixture of Weight-shared Heterogeneous Group Attention Experts for Dynamic Token-wise KV Optimization16 Jun 2025 0 repositories listed
-
CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following14 Jun 2025 0 repositories listed
-
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval14 Jun 2025 0 repositories listed
-
AC/DC: LLM-based Audio Comprehension via Dialogue Continuation12 Jun 2025 0 repositories listed
-
Conversational Search: From Fundamentals to Frontiers in the LLM Era12 Jun 2025 0 repositories listed
-
HalLoc: Token-level Localization of Hallucinations for Vision Language Models12 Jun 2025 0 repositories listed
-
Magistral12 Jun 2025 0 repositories listed
-
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models11 Jun 2025 0 repositories listed
-
LLaVA-c: Continual Improved Visual Instruction Tuning10 Jun 2025 0 repositories listed
-
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being10 Jun 2025 0 repositories listed
-
Aligning Text, Images, and 3D Structure Token-by-Token9 Jun 2025 0 repositories listed
-
Video Unlearning via Low-Rank Refusal Vector9 Jun 2025 0 repositories listed
-
Audio-Aware Large Language Models as Judges for Speaking Styles6 Jun 2025 0 repositories listed
-
On the Mechanism of Reasoning Pattern Selection in Reinforcement Learning for Language Models5 Jun 2025 0 repositories listed
-
RELIC: Evaluating Compositional Instruction Following via Language Recognition5 Jun 2025 0 repositories listed
-
SeedEdit 3.0: Fast and High-Quality Generative Image Editing5 Jun 2025 0 repositories listed
-
Unleashing Hour-Scale Video Training for Long Video-Language Understanding5 Jun 2025 0 repositories listed
-
Robust Anti-Backdoor Instruction Tuning in LVLMs4 Jun 2025 0 repositories listed
-
MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching3 Jun 2025 0 repositories listed
-
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs2 Jun 2025 0 repositories listed
-
TIIF-Bench: How Does Your T2I Model Follow Your Instructions?2 Jun 2025 0 repositories listed
-
PersianMedQA: Language-Centric Evaluation of LLMs in the Persian Medical Domain30 May 2025 0 repositories listed
-
ARC: Argument Representation and Coverage Analysis for Zero-Shot Long Document Summarization with Instruction Following LLMs29 May 2025 0 repositories listed
-
ChartMind: A Comprehensive Benchmark for Complex Real-world Multimodal Chart Question Answering29 May 2025 0 repositories listed
-
Differential Information: An Information-Theoretic Perspective on Preference Optimization29 May 2025 0 repositories listed
-
Adaptive Detoxification: Safeguarding General Capabilities of LLMs through Toxicity-Aware Knowledge Editing28 May 2025 0 repositories listed
-
LaMDAgent: An Autonomous Framework for Post-Training Pipeline Optimization via LLM Agents28 May 2025 0 repositories listed
-
PartInstruct: Part-level Instruction Following for Fine-grained Robot Manipulation27 May 2025 0 repositories listed
-
Evaluating Robustness of Large Audio Language Models to Audio Injection: An Empirical Study26 May 2025 0 repositories listed
-
From Alignment to Advancement: Bootstrapping Audio-Language Alignment with Synthetic Data26 May 2025 0 repositories listed
-
Leveraging Importance Sampling to Detach Alignment Modules from Large Language Models26 May 2025 0 repositories listed
-
StyleAR: Customizing Multimodal Autoregressive Model for Style-Aligned Text-to-Image Generation26 May 2025 0 repositories listed
-
RECAST: Strengthening LLMs' Complex Instruction Following with Constraint-Verifiable Data25 May 2025 0 repositories listed
-
MIDB: Multilingual Instruction Data Booster for Enhancing Multilingual Instruction Synthesis23 May 2025 0 repositories listed
-
In-Context Watermarks for Large Language Models22 May 2025 0 repositories listed
-
ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models22 May 2025 0 repositories listed
-
22 May 2025 0 repositories listed Syntology 3 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
ToDi: Token-wise Distillation via Fine-Grained Divergence Control22 May 2025 0 repositories listed
-
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective21 May 2025 0 repositories listed
-
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management21 May 2025 0 repositories listed
-
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought21 May 2025 0 repositories listed
-
Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning21 May 2025 0 repositories listed
-
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy21 May 2025 0 repositories listed
-
DecIF: Improving Instruction-Following through Meta-Decomposition20 May 2025 0 repositories listed
-
Domain Adaptation of VLM for Soccer Video Understanding20 May 2025 0 repositories listed
-
Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels20 May 2025 0 repositories listed
-
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training20 May 2025 0 repositories listed
-
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers19 May 2025 0 repositories listed
-
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 202519 May 2025 0 repositories listed
-
Multi-Level Aware Preference Learning: Enhancing RLHF for Complex Multi-Instruction Tasks19 May 2025 0 repositories listed
-
Rethinking Predictive Modeling for LLM Routing: When Simple kNN Beats Complex Learned Routers19 May 2025 0 repositories listed
-
CompBench: Benchmarking Complex Instruction-guided Image Editing18 May 2025 0 repositories listed
-
Enhancing Complex Instruction Following for Large Language Models with Mixture-of-Contexts Fine-tuning17 May 2025 0 repositories listed
-
When Thinking Fails: The Pitfalls of Reasoning for Instruction-Following in LLMs16 May 2025 0 repositories listed
-
HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages16 May 2025 0 repositories listed
-
Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Factor Mining16 May 2025 0 repositories listed
-
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation15 May 2025 0 repositories listed
-
Tests as Prompt: A Test-Driven-Development Benchmark for LLM Code Generation13 May 2025 0 repositories listed
-
Efficient Telecom Specific LLM: TSLAM-Mini with QLoRA and Digital Twin Data10 May 2025 0 repositories listed
-
Assessing Robustness to Spurious Correlations in Post-Training Language Models9 May 2025 0 repositories listed
-
T2VTextBench: A Human Evaluation Benchmark for Textual Control in Video Generation Models8 May 2025 0 repositories listed
-
Incentivizing Inclusive Contributions in Model Sharing Markets5 May 2025 0 repositories listed
-
PIPA: A Unified Evaluation Protocol for Diagnosing Interactive Planning Agents2 May 2025 0 repositories listed
-
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation1 May 2025 0 repositories listed
-
Ask, Fail, Repeat: Meeseeks, an Iterative Feedback Benchmark for LLMs' Multi-turn Instruction-Following Ability30 Apr 2025 0 repositories listed
-
UAV-VLN: End-to-End Vision Language guided Navigation for UAVs30 Apr 2025 0 repositories listed
-
CachePrune: Neural-Based Attribution Defense Against Indirect Prompt Injection Attacks29 Apr 2025 0 repositories listed
-
Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs24 Apr 2025 0 repositories listed
-
23 Apr 2025 0 repositories listed
-
ManipDreamer: Boosting Robotic Manipulation World Model with Action Tree and Visual Guidance23 Apr 2025 0 repositories listed
-
ParamΔ for Direct Weight Mixing: Post-Train Large Language Model at Zero Cost23 Apr 2025 0 repositories listed
-
DistilQwen2.5: Industrial Practices of Training Distilled Open Lightweight Language Models21 Apr 2025 0 repositories listed
-
Improving Instruct Models for Free: A Study on Partial Adaptation15 Apr 2025 0 repositories listed
-
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning12 Apr 2025 0 repositories listed
-
Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models10 Apr 2025 0 repositories listed
-
VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding10 Apr 2025 0 repositories listed
-
Holistic Capability Preservation: Towards Compact Yet Comprehensive Reasoning Models9 Apr 2025 0 repositories listed
-
Finding Fantastic Experts in MoEs: A Unified Study for Expert Dropping Strategies and Observations8 Apr 2025 0 repositories listed
-
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models8 Apr 2025 0 repositories listed
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.