{"url":"/method/tabert","slug":"tabert","name":"TaBERT","full_name":"TaBERT","full_name_withheld":false,"description_markdown":"**TaBERT** is a pretrained language model (LM) that jointly learns representations for natural language sentences and (semi-)structured tables. TaBERT is trained on a large corpus of 26 million tables and their English contexts. \r\n\r\nIn summary, TaBERT's process for learning representations for NL sentences is as follows: Given an utterance $u$ and a table $T$, TaBERT first creates a content snapshot of $T$. This snapshot consists of sampled rows that summarize the information in $T$ most relevant to the input utterance. The model then linearizes each row in the snapshot, concatenates each linearized row with the utterance, and uses the concatenated string as input to a Transformer model, which outputs row-wise encoding vectors of utterance tokens and cells. The encodings for all the rows in the snapshot are fed into a series of vertical self-attention layers, where a cell representation (or an utterance token representation) is computed by attending to vertically-aligned vectors of the same column (or the same NL token). Finally, representations for each utterance token and column are generated from a pooling layer.","description_state":"present","introduced_year":null,"introduced_by":{"title":"TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data","paper":"/paper/tabert-pretraining-for-joint-understanding-of","first_author":"Pengcheng Yin","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/tabert-pretraining-for-joint-understanding-of"},"source":{"url":"https://arxiv.org/abs/2005.08314v1","title":"TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Deep Tabular Learning","url":"/methods/category/deep-tabular-learning","pwc_aliases":[]}],"n_papers_tagged":3,"archive_num_papers":3,"papers_newest_first":[{"paper":"/paper/ait-qa-question-answering-dataset-over","title":"AIT-QA: Question Answering Dataset over Complex Tables in the Airline Industry","date":"2021-06-24","arxiv_id":"2106.12944","n_code_links":1,"syntology":null},{"paper":null,"title":"Sattiy at SemEval-2021 Task 9: An Ensemble Solution for Statement Verification and Evidence Finding with Tables","date":"2021-04-21","arxiv_id":"2104.10366","n_code_links":0,"syntology":null},{"paper":"/paper/tabert-pretraining-for-joint-understanding-of","title":"TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data","date":"2020-05-17","arxiv_id":"2005.08314","n_code_links":1,"syntology":null}],"papers_shown":3,"tasks":[{"task":"/task/semantic-parsing","name":"Semantic Parsing","papers":3},{"task":"/task/question-answering","name":"Question Answering","papers":2},{"task":"/task/articles","name":"Articles","papers":1},{"task":"/task/form","name":"Form","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/text-to-sql","name":"Text to SQL","papers":1},{"task":"/task/text-to-sql","name":"Text-To-SQL","papers":1}],"tasks_shown":7,"n_tasks":7,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tabert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}