Home › Datasets › task › Sign Language Translation
Sign Language Translation datasets
archive 2025-07-28
17 datasets carry the task tag "Sign Language Translation" (the task itself: Sign Language Translation), ordered by the archive's paper count. Page 1 of 1: 17 shown of 17. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Sign Language Translation datasets 1–17 of 17
Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
114 papers · 2 benchmarks
WLASL (Word-Level American Sign Language)
WLASL is a large video dataset for Word-Level American Sign Language (ASL) recognition, which features 2,000 common different words in ASL.
66 papers · 3 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
MSASL is a real-life large-scale sign language data set comprising over 25,000 annotated videos.
9 papers · 1 benchmark
BosphorusSign22k is a benchmark dataset for vision-based user-independent isolated Sign Language Recognition (SLR).
8 papers · 0 benchmarks
Content4All is a collection of six open research datasets aimed at automatic sign language translation research.
3 papers · 0 benchmarks
Sign languages are the primary means of communication for a large number of people worldwide.
2 papers · 0 benchmarks
ASLG-PC12 (English-ASL Gloss Parallel Corpus 2012)
An artificial corpus built using grammatical dependencies rules due to the lack of resources for Sign Language.
1 paper · 1 benchmark
AzSLD (AzSLD - Azerbaijani Sign Language Dataset)
The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL).
1 paper · 0 benchmarks
A large-scale gloss-free sign language translation dataset with 1,985 hours of videos, approximately 86 times larger than the previous CSL-Daily dataset.
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
We provide separate training, development and test data.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.