Papers › LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR

LEGATO: Large-scale End-to-end Generalizable Approach to Typeset OMR

23 Jun 2025arXiv:2506.19065archive 2025-07-28

Guang Yang, Victoria Ebert, Nazif Tamer, Luiza Pozzobon, Noah A. Smith

We propose Legato, a new end-to-end transformer model for optical music recognition (OMR). Legato is the first large-scale pretrained OMR model capable of recognizing full-page or multi-page typeset music scores and the first to generate documents in ABC notation, a concise, human-readable format for symbolic music. Bringing together a pretrained vision encoder with an ABC decoder trained on a dataset of more than 214K images, our model exhibits the strong ability to generalize across various typeset scores. We conduct experiments on a range of datasets and demonstrate that our model achieves state-of-the-art performance. Given the lack of a standardized evaluation for end-to-end OMR, we comprehensively compare our model against the previous state of the art using a diverse set of metrics.

PaperPDFCode

Code

guang-yng/legato officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Decoder

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

ABCSET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections