Home › Datasets › task › Sign Language Translation

Sign Language Translation datasets

archive 2025-07-28

17 datasets carry the task tag "Sign Language Translation" (the task itself: Sign Language Translation), ordered by the archive's paper count. Page 1 of 1: 17 shown of 17. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Sign Language Translation datasets 1–17 of 17

Over a period of three years (2009 - 2011) the daily news and weather forecast airings of the German public tv-station PHOENIX featuring sign language interpretation have been recorded and the weather forecasts of a subset of 386 editions…
114 papers · 2 benchmarks
WLASL (Word-Level American Sign Language)
WLASL is a large video dataset for Word-Level American Sign Language (ASL) recognition, which features 2,000 common different words in ASL.
66 papers · 3 benchmarks
CSL-Daily (Chinese Sign Language Corpus) is a large-scale continuous SLT dataset.
63 papers · 2 benchmarks
How2Sign (A Large-scale Multimodal Dataset for Continuous American Sign Language)
The How2Sign is a multimodal and multiview continuous American Sign Language (ASL) dataset consisting of a parallel corpus of more than 80 hours of sign language videos and a set of corresponding modalities including speech, English…
44 papers · 3 benchmarks
Large-scale American Sign Language (ASL) - English dataset collected from online video sites (e.g., YouTube).
15 papers · 0 benchmarks
MSASL is a real-life large-scale sign language data set comprising over 25,000 annotated videos.
9 papers · 1 benchmark
BosphorusSign22k is a benchmark dataset for vision-based user-independent isolated Sign Language Recognition (SLR).
8 papers · 0 benchmarks
Content4All is a collection of six open research datasets aimed at automatic sign language translation research.
3 papers · 0 benchmarks
Sign languages are the primary means of communication for a large number of people worldwide.
2 papers · 0 benchmarks
ASLG-PC12 (English-ASL Gloss Parallel Corpus 2012)
An artificial corpus built using grammatical dependencies rules due to the lack of resources for Sign Language.
1 paper · 1 benchmark
AzSLD (AzSLD - Azerbaijani Sign Language Dataset)
The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL).
1 paper · 0 benchmarks
A large-scale gloss-free sign language translation dataset with 1,985 hours of videos, approximately 86 times larger than the previous CSL-Daily dataset.
1 paper · 0 benchmarks
DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news…
1 paper · 0 benchmarks
LSA-T (Lengua de Señas Argentina - Traducción)
LSA-T is the first continuous Argentinian Sign Language (LSA) dataset.
1 paper · 1 benchmark
LSFB Datasets (French Belgian Sign Language Datasets)
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur.
1 paper · 0 benchmarks
Mediapi-RGB is a bilingual corpus of French Sign Language (LSF) and written French in the form of subtitled videos, accompanied by complementary data (various representations, segmentation, vocabulary, etc.).
1 paper · 1 benchmark
We provide separate training, development and test data.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.