Different Changes Require Different Reasoning: Change-Type-Specialized Experts for Robust Change Captioning
Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images.
Papers with a repository link where Syntology ran at least one harvested code sample. Each page shows two streams, newest first within each, counted separately: papers newer than the archive snapshot come from Syntology's graph Syntology; the rest are archive rows archive 2025-07-28. The two are never added together.
Cards 31–45 of 5,105 graph papers newer than 2025-07-28 with a Syntology-ran sample; this feed shows the newest 150, newest arXiv id first. Dates and the abstract sentence are from arXiv's metadata (CC0) for 5,105 of 5,105.
Change captioning is the task of generating natural language descriptions that explain the changes between a pair of images.
Scientific knowledge about AI models is produced faster than the community can organize it.
Humans need to study only a handful of well-written textbooks to master a discipline and attempt its hardest problems.
Graph data across diverse domains can expose valuable relational information to unauthorized representation learning, creating a pressing need for protection against such misuse.
Existing research on object hallucination in multimodal large language models (MLLMs) predominantly attributes the problem to language priors such as over-reliance on textual co-occurrence statistics.
The development of long-context Large Language Models (LLMs) is constrained by the memory bandwidth bottleneck and quadratic complexity of the attention mechanism during decoding.
End-to-end autonomous driving models plan future trajectories from raw sensor input.
While face recognition systems are widely deployed, ensuring their demographic reliability and robustness under uncontrolled visual conditions remains a critical challenge.
Learning materials properties from scarce labels and unlabeled crystals is a central challenge for data-driven materials discovery.
Collecting natural-language referring expressions along with region annotations, such as masks or boxes, is a major bottleneck in visual grounding (VG), as annotators must write descriptions that distinguish target…
3D Gaussian Splatting (3DGS) enables photorealistic real-time novel view synthesis, yet placing a virtual camera to capture a desired frame remains largely manual.
Scientific reasoning remains challenging for open-source models, largely due to the lack of high-quality scientific reasoning data.
Volumetric video enables immersive free viewpoint rendering of dynamic real world scenes, yet existing methods struggle with long sequences and complex motions, often leading to temporal instability and visual artifacts.
Loop Engineering is emerging as a practice for organizing development work around coding agents.
Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich.
Cards 31–45 of 31,700 archive papers with a code link, filtered to those where Syntology ran at least one sample (the archive chip labels the rows, the Syntology chip labels the filter; the count is the archive rows that pass it); this feed shows the newest 150, archive date first (newest archive date 2025-07-17). Within a month, dated rows come first, then the 819 undated rows placed by the month in their arXiv id. 1 archive paper with neither a date nor an arXiv id cannot be placed and is not listed.
Continuous-time Consistency Models (CMs) promise efficient few-step generation but face significant challenges with training instability.
Large language models (LLMs) excel at logical and algorithmic reasoning, yet their emotional intelligence (EQ) still lags far behind their cognitive prowess.
Next token prediction paradigm has been prevailing for autoregressive models in the era of LLMs.
Prompt injection attacks pose a significant security threat to LLM-integrated applications.
Simultaneous understanding and 3D reconstruction plays an important role in developing end-to-end embodied intelligent systems.
Inference-time computation techniques, analogous to human System 2 Thinking, have recently become popular for improving model performances.
While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's…
Unified image restoration is a significantly challenging task in low-level vision.
Visual object tracking has gained promising progress in past decades.
The architecture of multimodal large language models (MLLMs) commonly connects a vision encoder, often based on CLIP-ViT, to a large language model.
Residual connection has been extensively studied and widely applied at the model architecture level.
Recent advances in reinforcement learning have shown that language models can develop sophisticated reasoning through training on tasks with verifiable rewards, but these approaches depend on human-curated…
Benefiting from the advances in large language models and cross-modal alignment, existing multimodal large language models have achieved prominent performance in image and short video understanding.
Dataset distillation (DD) condenses large datasets into compact yet informative substitutes, preserving performance comparable to the original dataset while reducing storage, transmission costs, and computational…
The feed is static: 10 pages of up to 15 cards per stream, rebuilt with the site. Older papers are reachable from task, dataset and method pages and from search. No repository stars are tracked and nothing here is ranked by popularity. Machine-readable twin: JSON.