Browse State-of-the-Art › Language Modelling › Papers, page 74
Language Modelling
Papers archive 2025-07-28
archive papers tagged: 17,610 · with a code link: 7,012 · where Syntology ran a sample: 2,428 (2,027 with a run with no instrument failure, 401 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (2,428 of 17,610 tagged: 2,027 with a run with no instrument failure, 401 where every run was a failure of Syntology's instrument)
Page 74 of 177: papers 7,301 to 7,400 of 17,610, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
BiomechGPT: Towards a Biomechanically Fluent Multimodal Foundation Model for Clinically Relevant Motion Tasks24 May 2025 0 repositories listed
-
Building a Functional Machine Translation Corpus for Kpelle24 May 2025 0 repositories listed
-
Disentangling Knowledge Representations for Large Language Model Editing24 May 2025 0 repositories listed
-
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models24 May 2025 0 repositories listed
-
Inference Compute-Optimal Video Vision Language Models24 May 2025 0 repositories listed
-
metaTextGrad: Automatically optimizing language model optimizers24 May 2025 0 repositories listed
-
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors24 May 2025 0 repositories listed
-
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing24 May 2025 0 repositories listed
-
Skip-Thinking: Chunk-wise Chain-of-Thought Distillation Enable Smaller Language Models to Reason Better and Faster24 May 2025 0 repositories listed
-
Synthesizing and Adapting Error Correction Data for Mobile Large Language Model Applications24 May 2025 0 repositories listed
-
ELDeR: Getting Efficient LLMs through Data-Driven Regularized Layer-wise Pruning23 May 2025 0 repositories listed
-
Large language model as user daily behavior data generator: balancing population diversity and individual personality23 May 2025 0 repositories listed
-
Multi-agent Systems for Misinformation Lifecycle : Detection, Correction And Source Identification23 May 2025 0 repositories listed
-
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache23 May 2025 0 repositories listed
-
Plan-R1: Safe and Feasible Trajectory Planning as Language Modeling23 May 2025 0 repositories listed
-
QwenLong-CPRS: Towards ∞-LLMs with Dynamic Context Optimization23 May 2025 0 repositories listed
-
Retrieval Augmented Generation-based Large Language Models for Bridging Transportation Cybersecurity Legal Knowledge Gaps23 May 2025 0 repositories listed
-
Selection Mechanisms for Sequence Modeling using Linear State Space Models23 May 2025 0 repositories listed
-
Simulating Macroeconomic Expectations using LLM Agents23 May 2025 0 repositories listed
-
SpectraLDS: Provable Distillation for Linear Dynamical Systems23 May 2025 0 repositories listed
-
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback23 May 2025 0 repositories listed
-
Attention with Trained Embeddings Provably Selects Important Tokens22 May 2025 0 repositories listed
-
Beyond Correlation: Towards Causal Large Language Model Agents in Biomedicine22 May 2025 0 repositories listed
-
Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks22 May 2025 0 repositories listed
-
CTRAP: Embedding Collapse Trap to Safeguard Large Language Models from Harmful Fine-Tuning22 May 2025 0 repositories listed
-
DeepRec: Towards a Deep Dive Into the Item Space with Large Language Model Based Recommendation22 May 2025 0 repositories listed
-
Edge-First Language Model Inference: Models, Metrics, and Tradeoffs22 May 2025 0 repositories listed
-
Evaluating Large Language Model with Knowledge Oriented Language Specific Simple Question Answering22 May 2025 0 repositories listed
-
Incentivizing Dual Process Thinking for Efficient Large Language Model Reasoning22 May 2025 0 repositories listed
-
Incremental Sequence Classification with Temporal Consistency22 May 2025 0 repositories listed
-
INFERENCEDYNAMICS: Efficient Routing Across LLMs through Structured Capability and Knowledge Profiling22 May 2025 0 repositories listed
-
Large Language Model-Empowered Interactive Load Forecasting22 May 2025 0 repositories listed
-
Latent Principle Discovery for Language Model Self-Improvement22 May 2025 0 repositories listed
-
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning22 May 2025 0 repositories listed
-
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models22 May 2025 0 repositories listed
-
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing22 May 2025 0 repositories listed
-
On Multilingual Encoder Language Model Compression for Low-Resource Languages22 May 2025 0 repositories listed
-
Plan and Budget: Effective and Efficient Test-Time Scaling on Large Language Model Reasoning22 May 2025 0 repositories listed
-
Small-to-Large Generalization: Data Influences Models Consistently Across Scale22 May 2025 0 repositories listed
-
TensorAR: Refinement is All You Need in Autoregressive Image Generation22 May 2025 0 repositories listed
-
Aligning Dialogue Agents with Global Feedback via Large Language Model Reward Decomposition21 May 2025 0 repositories listed
-
Any Large Language Model Can Be a Reliable Judge: Debiasing with a Reasoning-based Bias Detector21 May 2025 0 repositories listed
-
CP-LLM: Context and Pixel Aware Large Language Model for Video Quality Assessment21 May 2025 0 repositories listed
-
DEBATE, TRAIN, EVOLVE: Self Evolution of Language Model Reasoning21 May 2025 0 repositories listed
-
Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering21 May 2025 0 repositories listed
-
Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective21 May 2025 0 repositories listed
-
Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model21 May 2025 0 repositories listed
-
Ensembling Sparse Autoencoders21 May 2025 0 repositories listed
-
Forging Time Series with Language: A Large Language Model Approach to Synthetic Data Generation21 May 2025 0 repositories listed
-
Internal and External Impacts of Natural Language Processing Papers21 May 2025 0 repositories listed
-
Likelihood Variance as Text Importance for Resampling Texts to Map Language Models21 May 2025 0 repositories listed
-
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions21 May 2025 0 repositories listed
-
Mechanistic evaluation of Transformers and state space models21 May 2025 0 repositories listed
-
MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling21 May 2025 0 repositories listed
-
Revealing Language Model Trajectories via Kullback-Leibler Divergence21 May 2025 0 repositories listed
-
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information21 May 2025 0 repositories listed
-
Self-GIVE: Associative Thinking from Limited Structured Knowledge for Enhanced Large Language Model Reasoning21 May 2025 0 repositories listed
-
Short-Range Dependency Effects on Transformer Instability and a Decomposed Attention Solution21 May 2025 0 repositories listed
-
Your Language Model Can Secretly Write Like Humans: Contrastive Paraphrase Attacks on LLM-Generated Text Detectors21 May 2025 0 repositories listed
-
Automated Journalistic Questions: A New Method for Extracting 5W1H in French20 May 2025 0 repositories listed
-
CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring20 May 2025 0 repositories listed
-
CtrlDiff: Boosting Large Diffusion Language Models with Dynamic Block Prediction and Controllable Generation20 May 2025 0 repositories listed
-
FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation20 May 2025 0 repositories listed
-
HausaNLP: Current Status, Challenges and Future Directions for Hausa Natural Language Processing20 May 2025 0 repositories listed
-
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising20 May 2025 0 repositories listed
-
Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications20 May 2025 0 repositories listed
-
MAS-KCL: Knowledge component graph structure learning with large language model-based agentic workflow20 May 2025 0 repositories listed
-
Structured Agent Distillation for Large Language Model20 May 2025 0 repositories listed
-
Studying the Role of Input-Neighbor Overlap in Retrieval-Augmented Language Models Training Efficiency20 May 2025 0 repositories listed
-
sudoLLM : On Multi-role Alignment of Language Models20 May 2025 0 repositories listed
-
TRATES: Trait-Specific Rubric-Assisted Cross-Prompt Essay Scoring20 May 2025 0 repositories listed
-
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation20 May 2025 0 repositories listed
-
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives20 May 2025 0 repositories listed
-
A*-Decoding: Token-Efficient Inference Scaling19 May 2025 0 repositories listed
-
A Physics-Inspired Optimizer: Velocity Regularized Adam19 May 2025 0 repositories listed
-
CMLFormer: A Dual Decoder Transformer with Switching Point Learning for Code-Mixed Language Modeling19 May 2025 0 repositories listed
-
Combining the Best of Both Worlds: A Method for Hybrid NMT and LLM Translation19 May 2025 0 repositories listed
-
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment19 May 2025 0 repositories listed
-
19 May 2025 0 repositories listed
-
On the Thinking-Language Modeling Gap in Large Language Models19 May 2025 0 repositories listed
-
ORQA: A Benchmark and Foundation Model for Holistic Operating Room Modeling19 May 2025 0 repositories listed
-
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model19 May 2025 0 repositories listed
-
ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model19 May 2025 0 repositories listed
-
Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping19 May 2025 0 repositories listed
-
SpatialLLM: From Multi-modality Data to Urban Spatial Intelligence19 May 2025 0 repositories listed
-
Structure-Aware Corpus Construction and User-Perception-Aligned Metrics for Large-Language-Model Code Completion19 May 2025 0 repositories listed
-
SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models19 May 2025 0 repositories listed
-
Tianyi: A Traditional Chinese Medicine all-rounder language model and its Real-World Clinical Practice19 May 2025 0 repositories listed
-
TinyAlign: Boosting Lightweight Vision-Language Models by Mitigating Modal Alignment Bottlenecks19 May 2025 0 repositories listed
-
VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection19 May 2025 0 repositories listed
-
VocalAgent: Large Language Models for Vocal Health Diagnostics with Safety-Aware Evaluation19 May 2025 0 repositories listed
-
Why Knowledge Distillation Works in Generative Models: A Minimal Working Explanation19 May 2025 0 repositories listed
-
Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent Systems18 May 2025 0 repositories listed
-
CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design18 May 2025 0 repositories listed
-
DS-ProGen: A Dual-Structure Deep Language Model for Functional Protein Design18 May 2025 0 repositories listed
-
From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling18 May 2025 0 repositories listed
-
LLM-Based User Simulation for Low-Knowledge Shilling Attacks on Recommender Systems18 May 2025 0 repositories listed
-
mCLM: A Function-Infused and Synthesis-Friendly Modular Chemical Language Model18 May 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.