Browse State-of-the-Art › Optical Character Recognition (OCR) › Papers, page 6
Optical Character Recognition (OCR)
Papers archive 2025-07-28
archive papers tagged: 1,243 · with a code link: 462 · where Syntology ran a sample: 76 (64 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (76 of 1,243 tagged: 64 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)
Page 6 of 13: papers 501 to 600 of 1,243, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Analyzing Patterns and Influence of Advertising in Print Newspapers16 May 2025 0 repositories listed
-
Object-Centric Representations Improve Policy Generalization in Robot Manipulation16 May 2025 0 repositories listed
-
A document processing pipeline for the construction of a dataset for topic modeling based on the judgments of the Italian Supreme Court13 May 2025 0 repositories listed
-
Gameplay Highlights Generation12 May 2025 0 repositories listed
-
Development of a WAZOBIA-Named Entity Recognition System10 May 2025 0 repositories listed
-
ChemRxivQuest: A Curated Chemistry Question-Answer Database Extracted from ChemRxiv Preprints8 May 2025 0 repositories listed
-
GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text Editing8 May 2025 0 repositories listed
-
Lost in OCR Translation? Vision-Based Approaches to Robust Document Retrieval8 May 2025 0 repositories listed
-
DOTA: Deformable Optimized Transformer Architecture for End-to-End Text Recognition with Retrieval-Augmented Generation7 May 2025 0 repositories listed
-
SymbioticRAG: Enhancing Document Intelligence Through Human-LLM Symbiotic Collaboration5 May 2025 0 repositories listed
-
Automated Parsing of Engineering Drawings for Structured Information Extraction Using a Fine-tuned Document Understanding Transformer2 May 2025 0 repositories listed
-
Entropy Heat-Mapping: Localizing GPT-Based OCR Errors with Sliding-Window Shannon Analysis30 Apr 2025 0 repositories listed
-
Guidelines for External Disturbance Factors in the Use of OCR in Real-World Environments21 Apr 2025 0 repositories listed
-
Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform21 Apr 2025 0 repositories listed
-
Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR15 Apr 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
Towards Calibration Enhanced Network by Inverse Adversarial Attack8 Apr 2025 0 repositories listed
-
Towards Visual Text Grounding of Multimodal Large Language Model7 Apr 2025 0 repositories listed
-
VISTA-OCR: Towards generative and interactive end to end OCR models4 Apr 2025 0 repositories listed
-
QID: Efficient Query-Informed ViTs in Data-Scarce Regimes for OCR-free Visual Document Understanding3 Apr 2025 0 repositories listed
-
Context-Independent OCR with Multimodal LLMs: Effects of Image Resolution and Visual Complexity31 Mar 2025 0 repositories listed
-
TFIC: End-to-End Text-Focused Image Compression for Coding for Machines25 Mar 2025 0 repositories listed
-
Slide2Text: Leveraging LLMs for Personalized Textbook Generation from PowerPoint Presentations22 Mar 2025 0 repositories listed
-
13 Mar 2025 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Revisiting Noise in Natural Language Processing for Computational Social Science10 Mar 2025 0 repositories listed
-
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model9 Mar 2025 0 repositories listed
-
AI-Driven Multi-Stage Computer Vision System for Defect Detection in Laser-Engraved Industrial Nameplates5 Mar 2025 0 repositories listed
-
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts25 Feb 2025 0 repositories listed
-
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI24 Feb 2025 0 repositories listed
-
Visual Zero-Shot E-Commerce Product Attribute Value Extraction21 Feb 2025 0 repositories listed
-
Harnessing PDF Data for Improving Japanese Large Multimodal Models20 Feb 2025 0 repositories listed
-
KITAB-Bench: A Comprehensive Multi-Domain Benchmark for Arabic OCR and Document Understanding20 Feb 2025 0 repositories listed
-
Corrupted but Not Broken: Understanding and Mitigating the Negative Impacts of Corrupted Data in Visual Instruction Tuning18 Feb 2025 0 repositories listed
-
Southern Newswire Corpus: A Large-Scale Dataset of Mid-Century Wire Articles Beyond the Front Page17 Feb 2025 0 repositories listed
-
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency13 Feb 2025 0 repositories listed
-
Visual Graph Question Answering with ASP and LLMs for Language Parsing13 Feb 2025 0 repositories listed
-
Adapting Multilingual Embedding Models to Historical Luxembourgish11 Feb 2025 0 repositories listed
-
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents6 Feb 2025 0 repositories listed
-
MME-Industry: A Cross-Industry Multimodal Evaluation Benchmark28 Jan 2025 0 repositories listed
-
LoCoML: A Framework for Real-World ML Inference Pipelines24 Jan 2025 0 repositories listed
-
Exploring AI-based System Design for Pixel-level Protected Health Information Detection in Medical Images16 Jan 2025 0 repositories listed
-
MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents15 Jan 2025 0 repositories listed
-
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations6 Jan 2025 0 repositories listed
-
SceneVTG++: Controllable Multilingual Visual Text Generation in the Wild6 Jan 2025 0 repositories listed
-
Emergency-Brake Simplex: Toward A Verifiably Safe Control-CPS Architecture for Abrupt Runtime Reachability Constraint Changes3 Jan 2025 0 repositories listed
-
Embedding Similarity Guided License Plate Super Resolution2 Jan 2025 0 repositories listed
-
CLIP is Almost All You Need: Towards Parameter-Efficient Scene Text Retrieval without OCR1 Jan 2025 0 repositories listed
-
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark1 Jan 2025 0 repositories listed
-
Optical Character Recognition using Convolutional Neural Networks for Ashokan Brahmi Inscriptions29 Dec 2024 0 repositories listed
-
ERPA: Efficient RPA Model Integrating OCR and LLMs for Intelligent Document Processing24 Dec 2024 0 repositories listed
-
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images24 Dec 2024 0 repositories listed
-
Leveraging Deep Learning with Multi-Head Attention for Accurate Extraction of Medicine from Handwritten Prescriptions24 Dec 2024 0 repositories listed
-
VORTEX: A Spatial Computing Framework for Optimized Drone Telemetry Extraction from First-Person View Flight Data24 Dec 2024 0 repositories listed
-
TextSleuth: Towards Explainable Tampered Text Detection19 Dec 2024 0 repositories listed
-
17 Dec 2024 0 repositories listed
-
Advanced ingestion process powered by LLM parsing for RAG system16 Dec 2024 0 repositories listed
-
Advancing Vehicle Plate Recognition: Multitasking Visual Language Models with VehiclePaliGemma14 Dec 2024 0 repositories listed
-
Enhancement of text recognition for hanja handwritten documents of Ancient Korea14 Dec 2024 0 repositories listed
-
One Filter to Deploy Them All: Robust Safety for Quadrupedal Navigation in Unknown Environments13 Dec 2024 0 repositories listed
-
AI Adoption to Combat Financial Crime: Study on Natural Language Processing in Adverse Media Screening of Financial Services in English and Bangla multilingual interpretation12 Dec 2024 0 repositories listed
-
DocSum: Domain-Adaptive Pre-training for Document Abstractive Summarization11 Dec 2024 0 repositories listed
-
DocVLM: Make Your VLM an Efficient Reader11 Dec 2024 0 repositories listed
-
POINTS1.5: Building a Vision-Language Model towards Real World Applications11 Dec 2024 0 repositories listed
-
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models6 Dec 2024 0 repositories listed
-
Text Change Detection in Multilingual Documents Using Image Comparison5 Dec 2024 0 repositories listed
-
CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy3 Dec 2024 0 repositories listed
-
Patchfinder: Leveraging Visual Language Models for Accurate Information Retrieval using Model Uncertainty3 Dec 2024 0 repositories listed
-
Arabic Handwritten Document OCR Solution with Binarization and Adaptive Scale Fusion Detection2 Dec 2024 0 repositories listed
-
AI-assisted summary of suicide risk Formulation29 Nov 2024 0 repositories listed
-
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs28 Nov 2024 0 repositories listed
-
VARCO-VISION: Expanding Frontiers in Korean Vision-Language Models28 Nov 2024 0 repositories listed
-
Towards Accessible Learning: Deep Learning-Based Potential Dysgraphia Detection and OCR for Potentially Dysgraphic Handwriting18 Nov 2024 0 repositories listed
-
Is Cognition consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding12 Nov 2024 0 repositories listed
-
Veri-Car: Towards Open-world Vehicle Information Retrieval11 Nov 2024 0 repositories listed
-
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding8 Nov 2024 0 repositories listed
-
NeKo: Toward Post Recognition Generative Correction Large Language Models with Task-Oriented Experts8 Nov 2024 0 repositories listed
-
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding7 Nov 2024 0 repositories listed
-
TAP-VL: Text Layout-Aware Pre-training for Enriched Vision-Language Models7 Nov 2024 0 repositories listed
-
Out-of-Distribution Recovery with Object-Centric Keypoint Inverse Policy for Visuomotor Imitation Learning5 Nov 2024 0 repositories listed
-
HIP: Hierarchical Point Modeling and Pre-training for Visual Information Extraction2 Nov 2024 0 repositories listed
-
Handwriting Recognition in Historical Documents with Multimodal LLM31 Oct 2024 0 repositories listed
-
Structured Analysis and Comparison of Alphabets in Historical Handwritten Ciphers29 Oct 2024 0 repositories listed
-
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding25 Oct 2024 0 repositories listed
-
Towards Visual Text Design Transfer Across Languages24 Oct 2024 0 repositories listed
-
Harnessing Webpage UIs for Text-Rich Visual Understanding17 Oct 2024 0 repositories listed
-
Reference-Based Post-OCR Processing with LLM for Diacritic Languages17 Oct 2024 0 repositories listed
-
Comparison of Image Preprocessing Techniques for Vehicle License Plate Recognition Using OCR: Performance and Accuracy Evaluation15 Oct 2024 0 repositories listed
-
ReLayout: Towards Real-World Document Understanding via Layout-enhanced Pre-training14 Oct 2024 0 repositories listed
-
ChartKG: A Knowledge-Graph-Based Representation for Chart Images13 Oct 2024 0 repositories listed
-
MIRAGE: Multimodal Identification and Recognition of Annotations in Indian General Prescriptions13 Oct 2024 0 repositories listed
-
TextMaster: Universal Controllable Text Edit13 Oct 2024 0 repositories listed
-
Unraveling Movie Genres through Cross-Attention Fusion of Bi-Modal Synergy of Poster12 Oct 2024 0 repositories listed
-
Automated Quality Control System for Canned Tuna Production using Artificial Vision8 Oct 2024 0 repositories listed
-
Mero Nagarikta: Advanced Nepali Citizenship Data Extractor with Deep Learning-Powered Text Detection and OCR8 Oct 2024 0 repositories listed
-
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends5 Oct 2024 0 repositories listed
-
Khattat: Enhancing Readability and Concept Representation of Semantic Typography1 Oct 2024 0 repositories listed
-
JaPOC: Japanese Post-OCR Correction Benchmark using Vouchers30 Sep 2024 0 repositories listed
-
30 Sep 2024 0 repositories listed
-
See then Tell: Enhancing Key Information Extraction with Vision Grounding29 Sep 2024 0 repositories listed
-
CodeSCAN: ScreenCast ANalysis for Video Programming Tutorials27 Sep 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.