{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/domain-adaptive-training-bert-for-response","title":"An Effective Domain Adaptive Post-Training Method for BERT in Response Selection","arxiv_id":"1908.04812","date":"2019-08-13","proceeding":null,"authors":["Taesun Whang","Dongyub Lee","Chanhee Lee","Kisu Yang","Dongsuk Oh","Heuiseok Lim"],"abstract":"We focus on multi-turn response selection in a retrieval-based dialog system. In this paper, we utilize the powerful pre-trained language model Bi-directional Encoder Representations from Transformer (BERT) for a multi-turn dialog system and propose a highly effective post-training method on domain-specific corpus. Although BERT is easily adopted to various NLP tasks and outperforms previous baselines of each task, it still has limitations if a task corpus is too focused on a certain domain. Post-training on domain-specific corpus (e.g., Ubuntu Corpus) helps the model to train contextualized representations and words that do not appear in general corpus (e.g., English Wikipedia). Experimental results show that our approach achieves new state-of-the-art on two response selection benchmarks (i.e., Ubuntu Corpus V1, Advising Corpus) performance improvement by 5.9% and 6% on R@1.","url_abs":"https://arxiv.org/abs/1908.04812v2","url_pdf":"https://arxiv.org/pdf/1908.04812v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"domain-adaptive-training-bert-for-response","repo_url":"https://github.com/taesunwhang/BERT-ResSel","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"conversational-response-selection","task_name":"Conversational Response Selection"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[{"method_slug":"absolute-position-encodings","method_name":"Absolute Position Encodings"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"attention","method_name":"Attention"},{"method_slug":"attention-dropout","method_name":"Attention Dropout"},{"method_slug":"bert","method_name":"BERT"},{"method_slug":"bpe","method_name":"BPE"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"label-smoothing","method_name":"Label Smoothing"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"linear-warmup-with-linear-decay","method_name":"Linear Warmup With Linear Decay"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"transformer","method_name":"Transformer"},{"method_slug":"weight-decay","method_name":"Weight Decay"},{"method_slug":"wordpiece","method_name":"WordPiece"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/conversational-response-selection-on-douban-1","task":"Conversational Response Selection","dataset":"Douban","model":"BERT","rank_in_archive_order":10,"of":16,"metrics":{"MAP":"0.591","MRR":"0.633","P@1":"0.454","R10@1":"0.280","R10@2":"0.470","R10@5":"0.828"},"uses_additional_data":false},{"leaderboard":"/sota/conversational-response-selection-on-rrs","task":"Conversational Response Selection","dataset":"RRS","model":"BERT","rank_in_archive_order":4,"of":7,"metrics":{"MAP":"0.625","MRR":"0.639","P@1":"0.453","R10@1":"0.404","R10@2":"0.606","R10@5":"0.875"},"uses_additional_data":false},{"leaderboard":"/sota/conversational-response-selection-on-rrs-1","task":"Conversational Response Selection","dataset":"RRS Ranking Test","model":"BERT","rank_in_archive_order":3,"of":4,"metrics":{"NDCG@3":"0.625","NDCG@5":"0.714"},"uses_additional_data":false},{"leaderboard":"/sota/conversational-response-selection-on-ubuntu-1","task":"Conversational Response Selection","dataset":"Ubuntu Dialogue (v1, Ranking)","model":"BERT-VFT","rank_in_archive_order":10,"of":25,"metrics":{"R10@1":"0.855","R10@2":"0.928","R10@5":"0.985"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1908.04812","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}