Hardware-Aware FP4 FlashAttention-4
Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink.
Papers with a repository link where Syntology ran at least one harvested code sample. Each page shows two streams, newest first within each, counted separately: papers newer than the archive snapshot come from Syntology's graph Syntology; the rest are archive rows archive 2025-07-28. The two are never added together.
Cards 16–30 of 5,105 graph papers newer than 2025-07-28 with a Syntology-ran sample; this feed shows the newest 150, newest arXiv id first. Dates and the abstract sentence are from arXiv's metadata (CC0) for 5,105 of 5,105.
Blackwell's 4-bit floating-point (FP4) tensor cores do not automatically make attention faster because softmax conversion and on-chip dependencies dominate once its matrix products shrink.
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery.
Mirrors are common in real-world images, yet producing geometrically consistent reflections with generative models remains challenging.
Understanding the composition of large-scale autonomous driving datasets is essential for safety, robustness, and reliable operation across domains.
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces.
Face recognition in unconstrained environments remains highly challenging due to diverse and extreme variations encountered in real-world scenarios.
Modern large language models (LLMs) rely on reinforcement learning to build strong capabilities in individual domains, but integrating those capabilities into a single deployable model remains challenging.
Test-Time Adaptation (TTA) has recently emerged as a promising strategy that allows the adaptation of pre-trained models to changing data distributions at deployment time, without access to any labels.
Text-Video Retrieval (TVR) retrieves videos that match a natural-language query, but extending image-text models such as CLIP to videos is fundamentally limited by the lack of temporal modeling.
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks.
Training Physics-Informed Neural Networks (PINNs) requires jointly optimizing physics residual and initial/boundary condition loss terms, which often induce conflicting gradients.
Dense self-supervised learning (SSL) is a powerful paradigm for learning without annotations the local descriptors required to solve dense medical imaging tasks.
Reinforcement learning is a natural way to post-train LLM agents for long-horizon interactive tasks judged only by end-of-task verification, yet a shared belief holds that outcome-only RL soon hits a ceiling on small…
Image retouching is commonly formulated as enhancing overall visual quality through color adjustment, but in practice, it also serves to emphasize visual focus by guiding viewers' attention toward a specific subject or…
Past literature on face recognition for monozygotic (("identical") twins points to facial marks and mirror asymmetry as possible directions for improved accuracy of twins recognition.
Cards 16–30 of 31,700 archive papers with a code link, filtered to those where Syntology ran at least one sample (the archive chip labels the rows, the Syntology chip labels the filter; the count is the archive rows that pass it); this feed shows the newest 150, archive date first (newest archive date 2025-07-17). Within a month, dated rows come first, then the 819 undated rows placed by the month in their arXiv id. 1 archive paper with neither a date nor an arXiv id cannot be placed and is not listed.
We introduce a full-stack framework that scales up reasoning in vision-language models (VLMs) to long videos, leveraging reinforcement learning.
In this paper, we investigate the distillation of time series reasoning capabilities into small, instruction-tuned language models as a step toward building interpretable time series foundation models.
We present a multi-agent system for automation of scientific research tasks, cmbagent.
Domain-Incremental Learning (DIL) focuses on continual learning in non-stationary environments, requiring models to adjust to evolving domains while preserving historical knowledge.
Large Language Models (LLMs) have achieved impressive accomplishments in recent years.
As language agents tackle increasingly complex tasks, they struggle with effective error correction and experience reuse across domains.
Sequence models like Transformers and RNNs often overallocate attention to irrelevant context, leading to noisy intermediate representations.
We introduce Skywork-R1V3, an advanced, open-source vision-language model (VLM) that pioneers a new approach to visual reasoning.
Notable breakthroughs in unified understanding and generation modeling have led to remarkable advancements in image understanding, reasoning, production and editing, yet current foundational models predominantly focus…
Graphical user interface (GUI) agents autonomously operate across platforms (e.g., Linux) to complete tasks by interacting with visual elements.
Vision-Language Models (VLMs) like CLIP have demonstrated remarkable generalization in zero- and few-shot settings, but adapting them efficiently to decentralized, heterogeneous data remains a challenge.
Historical documents represent an invaluable cultural heritage, yet have undergone significant degradation over time through tears, water erosion, and oxidation.
We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weights or activations.
Pre-trained vision-language models (VLMs) have advanced out-of-distribution (OOD) detection recently.
Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation.
The feed is static: 10 pages of up to 15 cards per stream, rebuilt with the site. Older papers are reachable from task, dataset and method pages and from search. No repository stars are tracked and nothing here is ranked by popularity. Machine-readable twin: JSON.