Best Practice Critic Optimization
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt.
Papers with a repository link where Syntology ran at least one harvested code sample. Each page shows two streams, newest first within each, counted separately: papers newer than the archive snapshot come from Syntology's graph Syntology; the rest are archive rows archive 2025-07-28. The two are never added together.
Cards 121–135 of 5,105 graph papers newer than 2025-07-28 with a Syntology-ran sample; this feed shows the newest 150, newest arXiv id first. Dates and the abstract sentence are from arXiv's metadata (CC0) for 5,105 of 5,105.
Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses for each prompt.
Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and active context.
Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs.
Narrow fine-tuning on small, domain-specific datasets can produce broad and surprising changes in model behavior-a phenomenon called weird generalization (WG).
Time series forecasting (TSF) is evolving toward multimodal and agentic settings, yet using foundation models remains uneconomical in resource-constrained scenarios, where compact, specialized forecasters are more…
Anomaly detection is often applied to data stored in relational databases, yet most existing methods require flattening multiple tables into a single feature matrix.
Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods.
The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities.
Accurate watch-time (WT) prediction is an important requirement for short-video recommendations.
Deep learning models for electrocardiogram (ECG) classification often suffer from significant performance degradation when deployed in unseen domains due to shifts in acquisition devices and patient populations.
Data assimilation (DA) is an essential tool for prediction and understanding in the geosciences.
Count data represented as a matrix of non-negative integer values, such as contingency tables, are prevalent across diverse domains.
As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation.
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal performance in other languages.
Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing unrelated metrics and citations to change even when the resulting PDF still builds.
Cards 121–135 of 31,700 archive papers with a code link, filtered to those where Syntology ran at least one sample (the archive chip labels the rows, the Syntology chip labels the filter; the count is the archive rows that pass it); this feed shows the newest 150, archive date first (newest archive date 2025-07-17). Within a month, dated rows come first, then the 819 undated rows placed by the month in their arXiv id. 1 archive paper with neither a date nor an arXiv id cannot be placed and is not listed.
Robust unlearning is crucial for safely deploying large language models (LLMs) in environments where data privacy, model safety, and regulatory compliance must be ensured.
Self-attention scales quadratically with input size, limiting its use for large-scale physical systems.
Self-supervised pre-trained audio networks have seen widespread adoption in real-world systems, particularly in multi-modal large language models.
Reinforcement learning (RL) with tree search has demonstrated superior performance in traditional reasoning tasks.
Self-supervised learning (SSL) has achieved major advances in natural images and video understanding, but challenges remain in domains like echocardiography (heart ultrasound) due to subtle anatomical structures,…
As the issue of global climate change becomes increasingly severe, the demand for research in climate science continues to grow.
Contrastive Language Audio Pretraining (CLAP) is a widely-used method to bridge the gap between audio and text domains.
Deep neural networks often learn and rely on spurious correlations, i.e., superficial associations between non-causal features and the targets.
Open multi-agent systems are increasingly important in modeling real-world applications, such as smart grids, swarm robotics, etc.
2D Gaussian Splatting (2DGS) has recently emerged as a promising method for novel view synthesis and surface reconstruction, offering better view-consistency and geometric accuracy than volumetric 3DGS.
In multimodal large language models (MLLMs), the length of input visual tokens is often significantly greater than that of their textual counterparts, leading to a high inference cost.
Obtaining high-quality labeled datasets is often costly, requiring either extensive human annotation or expensive experiments.
Uniform-state discrete diffusion models hold the promise of fast text generation due to their inherent ability to self-correct.
Understanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges.
Tabular in-context learning (ICL) has recently achieved state-of-the-art (SOTA) performance on several tabular prediction tasks.
The feed is static: 10 pages of up to 15 cards per stream, rebuilt with the site. Older papers are reachable from task, dataset and method pages and from search. No repository stars are tracked and nothing here is ranked by popularity. Machine-readable twin: JSON.