Datasets › SUC
SUC (Stockholm-Umeå Corpus)
The Stockholm-Umeå Corpus (SUC) is a collection of Swedish texts from the 1990s, consisting of one million words in total. The corpus is balanced, meaning that it contains various text types and stylistic levels. The texts are annotated with part-of-speech tags, morphological analysis, and lemma (all that can be considered gold standard data), as well as some structural and functional information.
Version 1.0 was developed in cooperation between Gunnel Källgren at Stockholm University and Eva Ejerhed at Umeå University and was made available in 1997 by the Department of Linguistics at Stockholm University.
Version 2.0 was made available in 2006 by Sofia Gustafsson Capkova and Britt Hartmann at the Department of Linguistics at Stockholm University. It contains the same texts as SUC 1.0 but is extended with some annotation. Additionally, SUC 2.0 contains bonus materials. TigerSUC is SUC 2.0 converted to TIGER-XML by Martin Volk. StorSUC is an additional SUC material of four million words.
Version 3.0 is available since 2012. It contains improved annotations and unannotated texts with seven million words.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- SUC
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections