{"url":"/dataset/avist","name":"AVisT","full_name":"A Benchmark for Visual Object Tracking in Adverse Visibility","description_markdown":"One of the key factors behind the recent success in visual tracking is the availability of dedicated benchmarks. While being greatly benefiting to the tracking research, existing benchmarks do not pose the same difficulty as before with recent trackers achieving higher performance mainly due to (i) the introduction of more sophisticated transformers-based methods and (ii) the lack of diverse scenarios with adverse visibility such as, severe weather conditions, camouflage and imaging effects.\r\nWe introduce AVisT, a dedicated benchmark for visual tracking in diverse scenarios with adverse visibility. AVisT comprises 120 challenging sequences with 80k annotated frames, spanning 18 diverse scenarios broadly grouped into five attributes with 42 object categories. The key contribution of AVisT is diverse and challenging scenarios covering severe weather conditions such as, dense fog, heavy rain and sandstorm; obstruction effects including, fire, sun glare and splashing water; adverse imaging effects such as, low-light; target effects including, small targets and distractor objects along with camouflage. We further benchmark 17 popular and recent trackers on AVisT with detailed analysis of their tracking performance across attributes, demonstrating a big room for improvement in performance. We believe that AVisT can greatly benefit the tracking community by complementing the existing benchmarks, in developing new creative tracking solutions in order to continue pushing the boundaries of the state-of-the-art. Our dataset along with the complete tracking performance evaluation is available.","description_withheld":null,"homepage":"https://arxiv.org/abs/2208.06888","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[{"name":"Visual Object Tracking","url":"/task/visual-object-tracking","datasets_with_task":"/datasets/task/visual-object-tracking"}],"languages":[],"variants":["AVisT"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-object-tracking-on-avist","task":"Visual Object Tracking","dataset_variant":"AVisT","rows":7,"metrics":["Success Rate"],"first_row_in_archive_order":{"model":"PiVOT-L","paper":"/paper/improving-visual-object-tracking-through","metrics":{"Success Rate":"62.2"},"code_links":[{"title":"chenshihfang/GOT","url":"https://github.com/chenshihfang/GOT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/improving-visual-object-tracking-through","title":"Improving Visual Object Tracking through Visual Prompting","date":"2024-09-27","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/unifying-visual-and-vision-language-tracking","title":"Unifying Visual and Vision-Language Tracking via Contrastive Learning","date":"2024-01-20","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/generalized-relation-modeling-for-transformer","title":"Generalized Relation Modeling for Transformer Tracking","date":"2023-03-29","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":0,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/transforming-model-prediction-for-tracking","title":"Transforming Model Prediction for Tracking","date":"2022-03-21","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/mixformer-end-to-end-tracking-with-iterative-1","title":"MixFormer: End-to-End Tracking with Iterative Mixed Attention","date":"2022-03-21","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/learning-spatio-temporal-transformer-for","title":"Learning Spatio-Temporal Transformer for Visual Tracking","date":"2021-03-31","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/2103-15436","title":"Transformer Tracking","date":"2021-03-29","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":1,"samples_unverified":2,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":5,"samples_harvested":8,"samples_ran":5,"samples_unverified":3,"pointer_only_for_licence":3,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}