A Physical Response-and-Memory Model for Muon Optimization
Training large language models is costly.
Papers with a repository link where Syntology ran at least one harvested code sample. Each page shows two streams, newest first within each, counted separately: papers newer than the archive snapshot come from Syntology's graph Syntology; the rest are archive rows archive 2025-07-28. The two are never added together.
Cards 136–150 of 5,105 graph papers newer than 2025-07-28 with a Syntology-ran sample; this feed shows the newest 150, newest arXiv id first. Dates and the abstract sentence are from arXiv's metadata (CC0) for 5,105 of 5,105.
Training large language models is costly.
Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decomposition and tool coordination but also accumulates self-generated text.
Multi-behavior recommendation (MBR) leverages auxiliary behavioral signals, such as clicks and add-to-cart, to enhance target behavior prediction like purchases.
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index…
When can an agent failure be caught?
Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality collapse, and incomplete representation of biochemical environments.
Generative segmentation provides an alternative to direct pixel-wise prediction by operating on learned latent representations, but effective image-to-mask translation must preserve target structure while remaining…
Decentralized federated learning (DFL) is a promising paradigm for autonomous nodes to collaboratively train AI models without relying on a central server.
Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic,…
A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort.
We introduce a new class of regression models for scan statistics on real-valued signals.
The components of a transformer communicate by writing to and reading from a shared residual stream, and mechanistic interpretability has mapped these connections by hand, one circuit at a time.
We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exploration process is dominated by stochastic variance.
LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software-engineering, and web tasks.
Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph information from visual observations.
Cards 136–150 of 31,700 archive papers with a code link, filtered to those where Syntology ran at least one sample (the archive chip labels the rows, the Syntology chip labels the filter; the count is the archive rows that pass it); this feed shows the newest 150, archive date first (newest archive date 2025-07-17). Within a month, dated rows come first, then the 819 undated rows placed by the month in their arXiv id. 1 archive paper with neither a date nor an arXiv id cannot be placed and is not listed.
This paper presents a novel method for analyzing the latent space geometry of generative models, including statistical physics models and diffusion models, by reconstructing the Fisher information metric.
Scene text retrieval has made significant progress with the assistance of accurate text localization.
We investigate how large language models respond to prompts that differ only in their token-level realization but preserve the same semantic intent, a phenomenon we call prompt variance.
The scale diversity of point cloud data presents significant challenges in developing unified representation learning techniques for 3D vision.
Recent work has identified retrieval heads (Wu et al., 2025b), a subset of attention heads responsible for retrieving salient information in long-context language models (LMs), as measured by their copy-paste behavior…
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing large language models (LLMs), with verification engineering playing a central role.
Cell instance segmentation is critical to analyzing biomedical images, yet accurately distinguishing tightly touching cells remains a persistent challenge.
Action segmentation is a core challenge in high-level video understanding, aiming to partition untrimmed videos into segments and assign each a label from a predefined action set.
We propose, implement, and compare with competitors a new architecture of equivariant neural networks based on geometric (Clifford) algebras: Generalized Lipschitz Group Equivariant Neural Networks (GLGENN).
Diffusion distillation is a widely used technique to reduce the sampling cost of diffusion models, yet it often requires extensive training, and the student performance tends to be degraded.
We propose Ming-Omni, a unified multimodal model capable of processing images, text, audio, and video, while demonstrating strong proficiency in both speech and image generation.
Classifier-free guidance (CFG) has become an essential component of modern diffusion models to enhance both generation quality and alignment with input conditions.
How well do AI systems perform in algorithm engineering for hard optimization problems in domains such as package-delivery routing, crew scheduling, factory production planning, and power-grid balancing?
Vision-Language models (VLMs) show impressive abilities to answer questions on visual inputs (e.g., counting objects in an image), yet demonstrate higher accuracies when performing an analogous task on text (e.g.,…
Existing acceleration techniques for video diffusion models often rely on uniform heuristics or time-embedding variants to skip timesteps and reuse cached features.
The feed is static: 10 pages of up to 15 cards per stream, rebuilt with the site. Older papers are reachable from task, dataset and method pages and from search. No repository stars are tracked and nothing here is ranked by popularity. Machine-readable twin: JSON.