Datasets › UD_Tagalog-NewsCrawl

UD_Tagalog-NewsCrawl

archive 2025-07-28

The Tagalog Universal Dependencies NewsCrawl dataset consists of annotated text extracted from the Leipzig Tagalog Corpus. Data included in the Leipzig Tagalog Corpus were crawled from Tagalog-language online news sites by the Leipzig University Institute for Computer Science.

The text data was automatically parsed and annotated by Angelina Aquino (University of the Philippines), and then manually corrected according the UD guidelines adapted for Tagalog by Elsie Marie Or (University of the Philippines), Maria Bardají Farré (University of Cologne), and Dr. Nikolaus Himmelmann (University of Cologne). Further verification and automated corrections were done by Lester James Miranda (Allen AI).

Due to the source of the data, several typos, grammatical errors, incomplete sentences, and Tagalog-English code-mixing can be found in the dataset.

Treebank structure

  • Train: 12495 sents, 286891 tokens
  • Dev: 1561 sents, 37045 tokens
  • Test: 1563 sents, 36974 tokens
Acknowledgments

Aside from the named persons in the previous section, the following also contributed to the project as manual annotators of the dataset:

  • Patricia Anne Asuncion
  • Paola Ellaine Luzon
  • Jenard Tricano
  • Mary Dianne Jamindang
  • Michael Wilson Rosero
  • Jim Bagano
  • Yeddah Joy Piedad
  • Farah Cunanan
  • Calen Manzano
  • Aien Gengania
  • Prince Heinreich Omang
  • Noah Cruz
  • Leila Ysabelle Suarez
  • Orlyn Joyce Esquivel
  • Andre Magpantay

The annotation project was made possible by the Deutsche Forschungsgemeinschaft (DFG)-funded project titled "Information distribution and language structure - correlation of grammatical expressions of the noun/verb distinction and lexical information content in Tagalog, Indonesian and German." The DFG project team is composed of Dr. Nikolaus Himmelmann and Maria Bardají Farré from the University of Cologne, and Dr. Gerhard Heyer, Dr. Michael Richter, and Tariq Yousef from the Leipzig University.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • UD_Tagalog-NewsCrawl

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections