{"url":"/dataset/azsld","name":"AzSLD","full_name":"AzSLD - Azerbaijani Sign Language Dataset","description_markdown":"The Azerbaijani Sign Language Dataset (AzSLD) is a comprehensive, large dataset designed to facilitate the development and evaluation of machine learning models for the recognition and translation of Azerbaijani Sign Language (AzSL). \r\n\r\nAzSLD is the first publicly available dataset focused on Azerbaijani Sign Language. It contributes to the global effort to improve accessibility for the deaf and hard-of-hearing community in Azerbaijan. The dataset aims to bridge the gap between technology and accessibility by providing high-quality data for researchers, developers, and practitioners working on sign language recognition or translation systems.\r\n\r\nThe data collection costs are covered by the \"Strengthening Data Analytics Research and Training Capacity through Establishment of dual Master of Science in Computer Science and Master of Science in Data Analytics (MSCS/DA) degree program at ADA University\" project, funded by BP and the Ministry of Education of the Republic of Azerbaijan.\r\n\r\nDataset Composition\r\nAzSLD is organized into three primary components:\r\n\r\n1. AzSLD_Sentences\r\nThis component contains video sequences of complete sentences in AzSL. It is designed to capture the fluidity and contextual nature of sign language, providing data for more complex language modeling tasks. It includes over 60 hours of high-definition video recordings, annotated with timestamped glosses for 500 distinct classes, enabling precise analysis and robust model training. Ground truth annotations of sentences for each class were added in a separate file. The videos were performed by 18 to 25 different signers, with a slight imbalance among them. \r\n\r\n2. AzSLD_Words\r\nThis component comprises a collection of short video samples representing frequently used words in Azerbaijani Sign Language. It is divided into two subsets:\r\n\r\nAzSLD_Words_100: Contains 100 commonly used words in AzSL.\r\nAzSLD_Words_200: Extends the first subset, including all 100 words from AzSLD_Words_100 along with an additional 100 words, for a total of 200 words.\r\nFolder names indicate the ground truth labels for the ease of word-level model evaluation.\r\n\r\n3. AzSLD_Fingerspelling\r\nThis component includes over 14,000 video and image samples of letters of the Azerbaijani alphabet. Each sign is captured from multiple angles to ensure comprehensive coverage of dactylology in AzSL. This component is ideal for tasks involving letter recognition and the integration of fingerspelling into broader sign language recognition systems.\r\n\r\nKey Features\r\nDouble-View Recordings\r\nThe dataset includes 10,104 synchronized video recordings from two camera angles to capture both frontal and side views of hand and body movements, ensuring that the subtle nuances of sign language are well-represented.\r\n\r\nDiverse Signers\r\nThe dataset features recordings from a diverse group of native AzSL signers, encompassing variations in age, gender, and signing style. This diversity is crucial for training models that are robust to variations in signing.\r\n\r\nDetailed Annotations\r\nEach video is annotated with comprehensive metadata, including the sign’s label (dactyl, word, or sentence), signer ID, and timestamped glosses for sentence-level signs. \r\n\r\nHigh-Quality Data Format\r\nThe dataset comprises RGB videos in high-definition (HD) resolution at 35 frames per second, accompanied by JSON files containing annotations and metadata. The data is systematically organized into folders by category for ease of navigation.\r\n\r\nEthical Transparency\r\n\r\nAll participants provided informed consent for collecting, publishing, and using the data, ensuring compliance with ethical research standards.\r\n\r\nAccessibility\r\n\r\nThe AzSLD is available under Creative Commons Attribution 4.0 International with free access for academic research through Zenodo.\r\n\r\nCitation: When using AzSLD in your research, please cite the following paper:\r\n\r\nAlishzade, N., Hasanov, J. (2024). AzSLD: Azerbaijani Sign Language Dataset for Dactyl, Word, and Sentence Translation with Baseline Software. [Journal Name], [Volume(Issue)], [Pages]. DOI: [DOI link].\r\n\r\nThe preprint is available at: https://arxiv.org/abs/2411.12865 \r\n\r\nContact:\r\nFor questions, feedback, or contributions, please contact the project team at: slr.project.ada@gmail.com","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.14222948","introduced_date":"2024-11-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/azsld-azerbaijani-sign-language-dataset-for","title":"AzSLD: Azerbaijani Sign Language Dataset for Fingerspelling, Word, and Sentence Translation with Baseline Software","first_author":"Nigar Alishzade","url":null},"license":{"name":"Creative Commons Attribution 4.0 International","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"RGB Video","url":"/datasets/modality/rgb-video"}],"tasks":[{"name":"Sign Language Recognition","url":"/task/sign-language-recognition","datasets_with_task":"/datasets/task/sign-language-recognition"},{"name":"Sign Language Translation","url":"/task/sign-language-translation","datasets_with_task":"/datasets/task/sign-language-translation"}],"languages":[{"name":"Azerbaijani","url":"/datasets/language/azerbaijani"}],"variants":["AzSLD"],"data_loaders":[{"repo":"https://github.com/ADA-SITE-JML/azsl_dataloader","url":"https://github.com/ADA-SITE-JML/azsl_dataloader","frameworks":["tf","pytorch"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}