{"url":"/method/lv-vit","slug":"lv-vit","name":"LV-ViT","full_name":"LV-ViT","full_name_withheld":false,"description_markdown":"**LV-ViT** is a type of [vision transformer](https://paperswithcode.com/method/vision-transformer) that uses token labelling as a training objective. Different from the standard training\r\nobjective of ViTs that computes the classification loss on an additional trainable class token, token labelling takes advantage of all the image patch tokens to compute the training loss in a dense manner. Specifically, token labeling reformulates the image classification problem into multiple token-level recognition problems and assigns each patch token with an individual location-specific supervision generated by a machine annotator.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2104.10858v3","title":"All Tokens Matter: Token Labeling for Training Better Vision Transformers","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Vision Transformers","url":"/methods/category/vision-transformers","pwc_aliases":["vision-transformer"]}],"n_papers_tagged":9,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/2408-01986","title":"DeMansia: Mamba Never Forgets Any Tokens","date":"2024-08-04","arxiv_id":"2408.01986","n_code_links":1,"syntology":null},{"paper":"/paper/multi-criteria-token-fusion-with-one-step","title":"Multi-criteria Token Fusion with One-step-ahead Attention for Efficient Vision Transformers","date":"2024-03-15","arxiv_id":"2403.10030","n_code_links":1,"syntology":{"ran":3,"of":5,"unverified":2,"pointer_only":0}},{"paper":null,"title":"TPC-ViT: Token Propagation Controller for Efficient Vision Transformer","date":"2024-01-03","arxiv_id":"2401.01470","n_code_links":0,"syntology":null},{"paper":"/paper/the-lottery-ticket-hypothesis-for-vision","title":"Data Level Lottery Ticket Hypothesis for Vision Transformers","date":"2022-11-02","arxiv_id":"2211.01484","n_code_links":1,"syntology":null},{"paper":"/paper/coarse-to-fine-vision-transformer","title":"CF-ViT: A General Coarse-to-Fine Method for Vision Transformer","date":"2022-03-08","arxiv_id":"2203.03821","n_code_links":1,"syntology":{"ran":1,"of":2,"unverified":1,"pointer_only":1}},{"paper":null,"title":"Make A Long Image Short: Adaptive Token Length for Vision Transformers","date":"2021-12-03","arxiv_id":"2112.01686","n_code_links":0,"syntology":null},{"paper":"/paper/self-slimmed-vision-transformer","title":"Self-slimmed Vision Transformer","date":"2021-11-24","arxiv_id":"2111.12624","n_code_links":1,"syntology":null},{"paper":null,"title":"Self-Slimming Vision Transformer","date":"2021-09-29","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/token-labeling-training-a-85-5-top-1-accuracy","title":"All Tokens Matter: Token Labeling for Training Better Vision Transformers","date":"2021-04-22","arxiv_id":"2104.10858","n_code_links":7,"syntology":{"ran":3,"of":5,"unverified":2,"pointer_only":3}}],"papers_shown":9,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":4},{"task":"/task/image-classification","name":"image-classification","papers":4},{"task":"/task/efficient-vits","name":"Efficient ViTs","papers":2},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":2},{"task":"/task/action-recognition-in-videos","name":"Action Recognition","papers":1},{"task":"/task/all","name":"All","papers":1},{"task":"/task/analogical-similarity","name":"Analogical Similarity","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/classification","name":"General Classification","papers":1},{"task":"/task/informativeness","name":"Informativeness","papers":1},{"task":"/task/mamba","name":"Mamba","papers":1},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1},{"task":"/task/state-space-models","name":"State Space Models","papers":1},{"task":"/task/token-reduction","name":"Token Reduction","papers":1}],"tasks_shown":14,"n_tasks":14,"usage_by_year":[{"year":"2021","papers":4},{"year":"2022","papers":2},{"year":"2024","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/lv-vit"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}