{"url":"/method/cross-view-training","slug":"cross-view-training","name":"Cross-View Training","full_name":"Cross-View Training","full_name_withheld":false,"description_markdown":"**Cross View Training**, or **CVT**, is a semi-supervised algorithm for training distributed word representations that makes use of unlabelled and labelled examples. \r\n\r\nCVT adds $k$ auxiliary prediction modules to the model, a Bi-[LSTM](https://paperswithcode.com/method/lstm) encoder, which are used when learning on unlabeled examples. A prediction module is usually a small neural network (e.g., a hidden layer followed by a [softmax](https://paperswithcode.com/method/softmax) layer). Each one takes as input an intermediate representation $h^j(x_i)$ produced by the model (e.g., the outputs of one of the LSTMs in a Bi-LSTM model). It outputs a distribution over labels $p\\_{j}^{\\theta}\\left(y\\mid{x\\_{i}}\\right)$.\r\n\r\nEach $h^j$ is chosen such that it only uses a part of the input $x_i$; the particular choice can depend on the task and model architecture. The auxiliary prediction modules are only used during training; the test-time prediction come from the primary prediction module that produces $p_\\theta$.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1809.08370v1","title":"Semi-Supervised Sequence Modeling with Cross-View Training","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Word Embeddings","url":"/methods/category/word-embeddings","pwc_aliases":[]}],"n_papers_tagged":8,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/cross-view-graph-consistency-learning-for","title":"Cross-View Graph Consistency Learning for Invariant Graph Representations","date":"2023-11-20","arxiv_id":"2311.11821","n_code_links":1,"syntology":null},{"paper":"/paper/an-efficient-self-supervised-cross-view","title":"An Efficient Self-Supervised Cross-View Training For Sentence Embedding","date":"2023-11-06","arxiv_id":"2311.03228","n_code_links":1,"syntology":null},{"paper":"/paper/camera-conditioned-stable-feature-generation","title":"Camera-Conditioned Stable Feature Generation for Isolated Camera Supervised Person Re-IDentification","date":"2022-03-29","arxiv_id":"2203.15210","n_code_links":1,"syntology":null},{"paper":null,"title":"Industry Scale Semi-Supervised Learning for Natural Language Understanding","date":"2021-03-29","arxiv_id":"2103.15871","n_code_links":0,"syntology":null},{"paper":null,"title":"To BERT or Not to BERT: Comparing Task-specific and Task-agnostic Semi-Supervised Approaches for Sequence Tagging","date":"2020-10-27","arxiv_id":"2010.14042","n_code_links":0,"syntology":null},{"paper":null,"title":"Semi-Supervised Semantic Role Labeling with Cross-View Training","date":"2019-11-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Semi-supervised Thai Sentence Segmentation Using Local and Distant Word Representations","date":"2019-08-04","arxiv_id":"1908.01294","n_code_links":0,"syntology":null},{"paper":"/paper/semi-supervised-sequence-modeling-with-cross","title":"Semi-Supervised Sequence Modeling with Cross-View Training","date":"2018-09-22","arxiv_id":"1809.08370","n_code_links":2,"syntology":null}],"papers_shown":8,"tasks":[{"task":"/task/sentence","name":"Sentence","papers":4},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/dependency-parsing","name":"Dependency Parsing","papers":2},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":2},{"task":"/task/attribute","name":"Attribute","papers":1},{"task":"/task/ccg-supertagging","name":"CCG Supertagging","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/feature-engineering","name":"Feature Engineering","papers":1},{"task":"/task/graph-representation-learning","name":"Graph Representation Learning","papers":1},{"task":"/task/intent-classification","name":"Intent Classification","papers":1},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/link-prediction","name":"Link Prediction","papers":1},{"task":"/task/machine-translation","name":"Machine Translation","papers":1},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":1},{"task":"/task/cg","name":"NER","papers":1},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1}],"tasks_shown":20,"n_tasks":34,"usage_by_year":[{"year":"2018","papers":1},{"year":"2019","papers":2},{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2022","papers":1},{"year":"2023","papers":2}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/cross-view-training"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}