Browse State-of-the-Art › Caption Generation › Papers, page 2
Caption Generation
Papers archive 2025-07-28
archive papers tagged: 310 · with a code link: 119 · where Syntology ran a sample: 33 (29 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (33 of 310 tagged: 29 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)
Page 2 of 4: papers 101 to 200 of 310, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
1 Dec 2020 1 repository listed
-
2 Aug 2020 1 repository listed
-
8 Jul 2020 1 repository listed
-
21 Jun 2020 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
27 Oct 2019 1 repository listed
-
10 Oct 2019 1 repository listed Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
10 Sep 2019 1 repository listed Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 3 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
27 Aug 2019 1 repository listed
-
25 May 2019 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
1 Apr 2019 1 repository listed
-
12 Oct 2018 1 repository listed
-
30 Mar 2018 1 repository listed
-
12 Mar 2018 1 repository listed
-
20 Jun 2017 1 repository listed
-
17 Aug 2016 1 repository listed
-
12 Aug 2016 1 repository listed
-
4 Apr 2016 1 repository listed
-
16 Sep 2015 1 repository listed
-
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning9 Jul 2025 0 repositories listed
-
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits11 Jun 2025 0 repositories listed
-
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation3 Jun 2025 0 repositories listed
-
NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID26 May 2025 0 repositories listed
-
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance25 May 2025 0 repositories listed
-
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks22 May 2025 0 repositories listed
-
Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives20 May 2025 0 repositories listed
-
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation24 Apr 2025 0 repositories listed
-
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training17 Apr 2025 0 repositories listed
-
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention3 Apr 2025 0 repositories listed
-
Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering29 Mar 2025 0 repositories listed
-
LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images20 Mar 2025 0 repositories listed
-
IDEA: Inverted Text with Cooperative Deformable Aggregation for Multi-modal Object Re-Identification13 Mar 2025 0 repositories listed
-
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models8 Mar 2025 0 repositories listed
-
Fine-Grained Video Captioning through Scene Graph Consolidation23 Feb 2025 0 repositories listed
-
LongCaptioning: Unlocking the Power of Long Caption Generation in Large Multimodal Models21 Feb 2025 0 repositories listed
-
Enhancing Chest X-ray Classification through Knowledge Injection in Cross-Modality Learning19 Feb 2025 0 repositories listed
-
FE-LWS: Refined Image-Text Representations via Decoder Stacking and Fused Encodings for Remote Sensing Image Captioning13 Feb 2025 0 repositories listed
-
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SCICAP Challenge 202331 Jan 2025 0 repositories listed
-
MAMS: Model-Agnostic Module Selection Framework for Video Captioning30 Jan 2025 0 repositories listed
-
Measuring and Mitigating Hallucinations in Vision-Language Dataset Generation for Remote Sensing24 Jan 2025 0 repositories listed
-
Understanding How Paper Writers Use AI-Generated Captions in Figure Caption Writing10 Jan 2025 0 repositories listed
-
Time Series Language Model for Descriptive Caption Generation3 Jan 2025 0 repositories listed
-
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning31 Dec 2024 0 repositories listed
-
Learning from Massive Human Videos for Universal Humanoid Pose Control18 Dec 2024 0 repositories listed
-
DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding2 Dec 2024 0 repositories listed
-
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains22 Nov 2024 0 repositories listed
-
Everything is a Video: Unifying Modalities through Next-Frame Prediction15 Nov 2024 0 repositories listed
-
Grounded Video Caption Generation12 Nov 2024 0 repositories listed
-
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning12 Oct 2024 0 repositories listed
-
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer17 Sep 2024 0 repositories listed
-
CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving19 Aug 2024 0 repositories listed
-
14 Aug 2024 0 repositories listed
-
13 Aug 2024 0 repositories listed
-
XMeCap: Meme Caption Generation with Sub-Image Adaptability24 Jul 2024 0 repositories listed
-
Explainable Image Captioning using CNN- CNN architecture and Hierarchical Attention28 Jun 2024 0 repositories listed
-
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?20 Jun 2024 0 repositories listed
-
DS@BioMed at ImageCLEFmedical Caption 2024: Enhanced Attention Mechanisms in Medical Caption Generation through Concept Detection Integration1 Jun 2024 0 repositories listed
-
Multi-Modal Generative Embedding Model29 May 2024 0 repositories listed
-
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation22 May 2024 0 repositories listed
-
MICap: A Unified Model for Identity-aware Movie Descriptions19 May 2024 0 repositories listed
-
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation30 Apr 2024 0 repositories listed
-
The Solution for the ICCV 2023 1st Scientific Figure Captioning Challenge26 Mar 2024 0 repositories listed
-
LuoJiaHOG: A Hierarchy Oriented Geo-aware Image Caption Dataset for Remote Sensing Image-Text Retrival16 Mar 2024 0 repositories listed
-
PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide Image Classification and Captioning13 Mar 2024 0 repositories listed
-
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback11 Mar 2024 0 repositories listed
-
LLMs in Political Science: Heralding a New Era of Visual Analysis29 Feb 2024 0 repositories listed
-
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation18 Jan 2024 0 repositories listed
-
Social Media Ready Caption Generation for Brands3 Jan 2024 0 repositories listed
-
BEV-TSR: Text-Scene Retrieval in BEV Space for Autonomous Driving2 Jan 2024 0 repositories listed
-
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning25 Dec 2023 0 repositories listed
-
Enhancing Image Captioning with Neural Models1 Dec 2023 0 repositories listed
-
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers27 Nov 2023 0 repositories listed
-
DECap: Towards Generalized Explicit Caption Editing via Diffusion Mechanism25 Nov 2023 0 repositories listed
-
Dense Video Captioning: A Survey of Techniques, Datasets and Evaluation Protocols5 Nov 2023 0 repositories listed
-
Visual Analytics for Efficient Image Exploration and User-Guided Image Captioning2 Nov 2023 0 repositories listed
-
LoHoRavens: A Long-Horizon Language-Conditioned Benchmark for Robotic Tabletop Manipulation18 Oct 2023 0 repositories listed
-
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools16 Oct 2023 0 repositories listed
-
A Comparative Study of Pre-trained CNNs and GRU-Based Attention for Image Caption Generation11 Oct 2023 0 repositories listed
-
FaceGemma: Enhancing Image Captioning with Facial Attributes for Portrait Images24 Sep 2023 0 repositories listed
-
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning20 Sep 2023 0 repositories listed
-
ViCo: Engaging Video Comment Generation with Human Preference Rewards22 Aug 2023 0 repositories listed
-
AIC-AB NET: A Neural Network for Image Captioning with Spatial Attention and Text Attributes14 Jul 2023 0 repositories listed
-
6 Jul 2023 0 repositories listed
-
Knowledge Distillation for Efficient Audio-Visual Video Captioning16 Jun 2023 0 repositories listed
-
CapText: Large Language Model-based Caption Generation From Image Context and Description1 Jun 2023 0 repositories listed
-
RealignDiff: Boosting Text-to-Image Diffusion Model with Coarse-to-fine Semantic Re-alignment31 May 2023 0 repositories listed
-
HAAV: Hierarchical Aggregation of Augmented Views for Image Captioning25 May 2023 0 repositories listed
-
DiffCap: Exploring Continuous Diffusion on Image Captioning20 May 2023 0 repositories listed
-
Efficient Audio Captioning Transformer with Patchout and Text Guidance6 Apr 2023 0 repositories listed
-
Taming Encoder for Zero Fine-tuning Image Customization with Text-to-Image Diffusion Models5 Apr 2023 0 repositories listed
-
Multi-modal reward for visual relationships-based image captioning19 Mar 2023 0 repositories listed
-
GNNFormer: A Graph-based Framework for Cytopathology Report Generation17 Mar 2023 0 repositories listed
-
Summaries as Captions: Generating Figure Captions for Scientific Documents with Automated Text Summarization23 Feb 2023 0 repositories listed
-
Stacked Cross-modal Feature Consolidation Attention Networks for Image Captioning8 Feb 2023 0 repositories listed
-
Uncertainty-Aware Image Captioning30 Nov 2022 0 repositories listed
-
22 Nov 2022 0 repositories listed
-
Image Caption Generation for Low-Resource Assamese Language1 Nov 2022 0 repositories listed
-
Generating image captions with external encyclopedic knowledge10 Oct 2022 0 repositories listed
-
REST: REtrieve & Self-Train for generative action recognition29 Sep 2022 0 repositories listed
-
Medical Image Captioning via Generative Pretrained Transformers28 Sep 2022 0 repositories listed
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.