{"url":"/sota/click-through-rate-prediction-on-criteo","task":{"name":"Click-Through Rate Prediction","url":"/task/click-through-rate-prediction","note":null},"dataset":{"name":"Criteo","url":"/dataset/criteo"},"category":"Miscellaneous","categories":["Miscellaneous"],"category_note":null,"description":"Click-through rate prediction is the task of predicting the likelihood that something on a website (such as an advertisement) will be clicked.\r\n\r\n<span style=\"color:grey; opacity: 0.6\">( Image credit: [Deep Spatio-Temporal Neural Networks for Click-Through Rate Prediction](https://arxiv.org/pdf/1906.03776v2.pdf) )</span>","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["AUC","Log Loss"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"AUC":"higher","Log Loss":"lower"}},"counts":{"rows":39,"rows_with_code":34,"rows_with_paper_page":39,"rows_dated":39,"rows_using_additional_data":0},"rows":[{"rank_in_archive_order":1,"model":"QNN-α","metrics":{"AUC":"0.8163","Log Loss":"0.4358"},"uses_additional_data":false,"paper_date":"2025-05-23","paper":"/paper/revisiting-feature-interactions-from-the","paper_url":"https://arxiv.org/abs/2505.17999v1","paper_title":"Revisiting Feature Interactions from the Perspective of Quadratic Neural Networks for Click-through Rate Prediction","code":"https://github.com/salmon1802/QNN","n_code_links":1,"syntology":null},{"rank_in_archive_order":2,"model":"FCN","metrics":{"AUC":"0.8162","Log Loss":"0.4358"},"uses_additional_data":false,"paper_date":"2024-07-18","paper":"/paper/dcnv3-towards-next-generation-deep-cross","paper_url":"https://arxiv.org/abs/2407.13349v7","paper_title":"FCN: Fusing Exponential and Linear Cross Network for Click-Through Rate Prediction","code":"https://github.com/reczoo/FuxiCTR","n_code_links":2,"syntology":null},{"rank_in_archive_order":3,"model":"GDCN","metrics":{"AUC":"0.8161","Log Loss":"0.4360"},"uses_additional_data":false,"paper_date":"2023-11-08","paper":"/paper/towards-deeper-lighter-and-interpretable-1","paper_url":"https://arxiv.org/abs/2311.04635v1","paper_title":"Towards Deeper, Lighter and Interpretable Cross Network for CTR Prediction","code":"https://github.com/xue-pai/FuxiCTR","n_code_links":3,"syntology":{"n_ran":0,"n_unverified":5,"n_samples":5,"n_pointer_only_licence":5}},{"rank_in_archive_order":4,"model":"MemoNet","metrics":{"AUC":"0.8152"},"uses_additional_data":false,"paper_date":"2022-10-25","paper":"/paper/memonet-memorizing-representations-of-all","paper_url":"https://arxiv.org/abs/2211.01334v3","paper_title":"MemoNet: Memorizing All Cross Features' Representations Efficiently via Multi-Hash Codebook Network for CTR Prediction","code":"https://github.com/ptzhangAlg/RecAlg","n_code_links":1,"syntology":null},{"rank_in_archive_order":5,"model":"TF4CTR","metrics":{"AUC":"0.8150"},"uses_additional_data":false,"paper_date":"2024-05-06","paper":"/paper/tf4ctr-twin-focus-framework-for-ctr","paper_url":"https://arxiv.org/abs/2405.03167v3","paper_title":"TF4CTR: Twin Focus Framework for CTR Prediction via Adaptive Sample Differentiation","code":"https://github.com/salmon1802/tf4ctr","n_code_links":1,"syntology":null},{"rank_in_archive_order":6,"model":"FinalMLP + MMBAttn","metrics":{"AUC":"0.81497"},"uses_additional_data":false,"paper_date":"2023-08-25","paper":"/paper/mmbattn-max-mean-and-bit-wise-attention-for","paper_url":"https://arxiv.org/abs/2308.13187v1","paper_title":"MMBAttn: Max-Mean and Bit-wise Attention for CTR Prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":7,"model":"FinalMLP","metrics":{"AUC":"0.8149"},"uses_additional_data":false,"paper_date":"2023-04-03","paper":"/paper/finalmlp-an-enhanced-two-stream-mlp-model-for-1","paper_url":"https://arxiv.org/abs/2304.00902v4","paper_title":"FinalMLP: An Enhanced Two-Stream MLP Model for CTR Prediction","code":"https://github.com/reczoo/FuxiCTR","n_code_links":4,"syntology":{"n_ran":13,"n_unverified":3,"n_samples":16,"n_pointer_only_licence":0}},{"rank_in_archive_order":8,"model":"CETN","metrics":{"AUC":"0.8148","Log Loss":"0.4373"},"uses_additional_data":false,"paper_date":"2023-12-15","paper":"/paper/cetn-contrast-enhanced-through-network-for","paper_url":"https://arxiv.org/abs/2312.09715v2","paper_title":"CETN: Contrast-enhanced Through Network for CTR Prediction","code":"https://github.com/salmon1802/cetn","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"STEC","metrics":{"AUC":"0.8143","Log Loss":"0.4379"},"uses_additional_data":false,"paper_date":"2023-08-29","paper":"/paper/stec-see-through-transformer-based-encoder","paper_url":"https://arxiv.org/abs/2308.15033v2","paper_title":"STEC: See-Through Transformer-based Encoder for CTR Prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":10,"model":"DNN + MMBAttn","metrics":{"AUC":"0.8143"},"uses_additional_data":false,"paper_date":"2023-08-25","paper":"/paper/mmbattn-max-mean-and-bit-wise-attention-for","paper_url":"https://arxiv.org/abs/2308.13187v1","paper_title":"MMBAttn: Max-Mean and Bit-wise Attention for CTR Prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":11,"model":"MaskNet","metrics":{"AUC":"0.8131"},"uses_additional_data":false,"paper_date":"2021-02-09","paper":"/paper/masknet-introducing-feature-wise","paper_url":"https://arxiv.org/abs/2102.07619v2","paper_title":"MaskNet: Introducing Feature-Wise Multiplication to CTR Ranking Models by Instance-Guided Mask","code":"https://github.com/twitter/the-algorithm","n_code_links":21,"syntology":null},{"rank_in_archive_order":12,"model":"DeepLight","metrics":{"AUC":"0.8123","Log Loss":"0.4395"},"uses_additional_data":false,"paper_date":"2020-02-17","paper":"/paper/a-sparse-deep-factorization-machine-for","paper_url":"https://arxiv.org/abs/2002.06987v3","paper_title":"DeepLight: Deep Lightweight Feature Interactions for Accelerating CTR Predictions in Ad Serving","code":"https://github.com/WayneDW/DeepLight_Deep-Lightweight-Feature-Interactions","n_code_links":2,"syntology":null},{"rank_in_archive_order":13,"model":"CELS","metrics":{"AUC":"0.8117","Log Loss":"0.4400"},"uses_additional_data":false,"paper_date":"2023-08-01","paper":"/paper/cognitive-evolutionary-search-to-select","paper_url":"https://dl.acm.org/doi/10.1145/3580305.3599277","paper_title":"Cognitive Evolutionary Search to Select Feature Interactions for Click-Through Rate Prediction","code":"https://github.com/RunlongYu/CELS","n_code_links":2,"syntology":null},{"rank_in_archive_order":14,"model":"OptFS","metrics":{"AUC":"0.8116","Log Loss":"0.4401"},"uses_additional_data":false,"paper_date":"2023-01-26","paper":"/paper/optimizing-feature-set-for-click-through-rate","paper_url":"https://arxiv.org/abs/2301.10909v2","paper_title":"Optimizing Feature Set for Click-Through Rate Prediction","code":"https://github.com/fuyuanlyu/optfs","n_code_links":1,"syntology":{"n_ran":1,"n_unverified":4,"n_samples":5,"n_pointer_only_licence":0}},{"rank_in_archive_order":15,"model":"DCN V2","metrics":{"AUC":"0.8115","Log Loss":"0.4406"},"uses_additional_data":false,"paper_date":"2020-08-19","paper":"/paper/dcn-m-improved-deep-cross-network-for-feature","paper_url":"https://arxiv.org/abs/2008.13535v2","paper_title":"DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems","code":"https://github.com/tensorflow/models/tree/master/official/recommendation/ranking","n_code_links":12,"syntology":{"n_ran":2,"n_unverified":12,"n_samples":14,"n_pointer_only_licence":0}},{"rank_in_archive_order":16,"model":"OptEmbed","metrics":{"AUC":"0.8114","Log Loss":"0.44"},"uses_additional_data":false,"paper_date":"2022-08-09","paper":"/paper/optembed-learning-optimal-embedding-table-for","paper_url":"https://arxiv.org/abs/2208.04482v2","paper_title":"OptEmbed: Learning Optimal Embedding Table for Click-through Rate Prediction","code":"https://github.com/fuyuanlyu/optembed","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":1,"n_samples":1,"n_pointer_only_licence":0}},{"rank_in_archive_order":17,"model":"ContextNet","metrics":{"AUC":"0.8113"},"uses_additional_data":false,"paper_date":"2021-07-26","paper":"/paper/contextnet-a-click-through-rate-prediction","paper_url":"https://arxiv.org/abs/2107.12025v1","paper_title":"ContextNet: A Click-Through Rate Prediction Framework Using Contextual information to Refine Feature Embedding","code":"https://github.com/QunBB/DeepLearning/blob/main/recommendation/rank/contextnet.py","n_code_links":4,"syntology":null},{"rank_in_archive_order":18,"model":"FiBiNet++","metrics":{"AUC":"0.8110"},"uses_additional_data":false,"paper_date":"2022-09-12","paper":"/paper/fibinet-improving-fibinet-by-greatly-reducing","paper_url":"https://arxiv.org/abs/2209.05016v2","paper_title":"FiBiNet++: Reducing Model Size by Low Rank Feature Interaction Layer for CTR Prediction","code":"https://github.com/alibaba/EasyRec/blob/master/docs/source/models/fibinet.md","n_code_links":5,"syntology":null},{"rank_in_archive_order":19,"model":"NormDNN","metrics":{"AUC":"0.8107"},"uses_additional_data":false,"paper_date":"2020-06-23","paper":"/paper/correct-normalization-matters-understanding","paper_url":"https://arxiv.org/abs/2006.12753v2","paper_title":"Correct Normalization Matters: Understanding the Effect of Normalization On Deep Neural Network Models For Click-Through Rate Prediction","code":"https://github.com/alibaba/EasyRec/blob/master/easy_rec/python/input/criteo_input.py","n_code_links":1,"syntology":null},{"rank_in_archive_order":20,"model":"DeepFFM","metrics":{"AUC":"0.8104","Log Loss":"0.4416"},"uses_additional_data":false,"paper_date":"2019-05-15","paper":"/paper/fat-deepffm-field-attentive-deep-field-aware","paper_url":"https://arxiv.org/abs/1905.06336v1","paper_title":"FAT-DeepFFM: Field Attentive Deep Field-aware Factorization Machine","code":"https://github.com/PaddlePaddle/PaddleRec/tree/master/models/rank/fat_deepffm","n_code_links":12,"syntology":null},{"rank_in_archive_order":21,"model":"FiBiNET","metrics":{"AUC":"0.8103","Log Loss":"0.4423"},"uses_additional_data":false,"paper_date":"2019-05-23","paper":"/paper/fibinet-combining-feature-importance-and","paper_url":"https://arxiv.org/abs/1905.09433v1","paper_title":"FiBiNET: Combining Feature Importance and Bilinear feature Interaction for Click-Through Rate Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":31,"syntology":{"n_ran":2,"n_unverified":24,"n_samples":26,"n_pointer_only_licence":3}},{"rank_in_archive_order":22,"model":"OptInter","metrics":{"AUC":"0.8101","Log Loss":"0.4417"},"uses_additional_data":false,"paper_date":"2021-08-03","paper":"/paper/memorize-factorize-or-be-naive-learning","paper_url":"https://arxiv.org/abs/2108.01265v3","paper_title":"Memorize, Factorize, or be Naïve: Learning Optimal Feature Interaction Methods for CTR Prediction","code":"https://github.com/fuyuanlyu/OptInter","n_code_links":1,"syntology":{"n_ran":0,"n_unverified":9,"n_samples":9,"n_pointer_only_licence":0}},{"rank_in_archive_order":23,"model":"GateNet","metrics":{"AUC":"0.8100"},"uses_additional_data":false,"paper_date":"2020-07-06","paper":"/paper/gatenet-gating-enhanced-deep-network-for","paper_url":"https://arxiv.org/abs/2007.03519v1","paper_title":"GateNet: Gating-Enhanced Deep Network for Click-Through Rate Prediction","code":"https://github.com/PaddlePaddle/PaddleRec/tree/master/models/rank/gatenet","n_code_links":5,"syntology":null},{"rank_in_archive_order":24,"model":"AFN+","metrics":{"AUC":"0.8074"},"uses_additional_data":false,"paper_date":"2019-09-07","paper":"/paper/adaptive-factorization-network-learning","paper_url":"https://arxiv.org/abs/1909.03276v2","paper_title":"Adaptive Factorization Network: Learning Adaptive-Order Feature Interactions","code":"https://github.com/shenweichen/DeepCTR-Torch","n_code_links":4,"syntology":{"n_ran":0,"n_unverified":2,"n_samples":2,"n_pointer_only_licence":0}},{"rank_in_archive_order":25,"model":"XCrossNet","metrics":{"AUC":"0.8067"},"uses_additional_data":false,"paper_date":"2021-04-22","paper":"/paper/xcrossnet-feature-structure-oriented-learning","paper_url":"https://arxiv.org/abs/2104.10907v1","paper_title":"XCrossNet: Feature Structure-Oriented Learning for Click-Through Rate Prediction","code":"https://github.com/bigdata-ustc/XCrossNet","n_code_links":1,"syntology":null},{"rank_in_archive_order":26,"model":"Fi-GNN","metrics":{"AUC":"0.8062","Log Loss":"0.4453"},"uses_additional_data":false,"paper_date":"2019-10-12","paper":"/paper/fi-gnn-modeling-feature-interactions-via","paper_url":"https://arxiv.org/abs/1910.05552v2","paper_title":"Fi-GNN: Modeling Feature Interactions via Graph Neural Networks for CTR Prediction","code":"https://github.com/xue-pai/FuxiCTR","n_code_links":5,"syntology":null},{"rank_in_archive_order":27,"model":"AutoInt","metrics":{"AUC":"0.8061","Log Loss":"0.4454"},"uses_additional_data":false,"paper_date":"2018-10-29","paper":"/paper/autoint-automatic-feature-interaction","paper_url":"https://arxiv.org/abs/1810.11921v2","paper_title":"AutoInt: Automatic Feature Interaction Learning via Self-Attentive Neural Networks","code":"https://github.com/shenweichen/DeepCTR","n_code_links":19,"syntology":{"n_ran":0,"n_unverified":3,"n_samples":3,"n_pointer_only_licence":0}},{"rank_in_archive_order":28,"model":"Clustered Compositional Embeddings","metrics":{"AUC":"0.806","Log Loss":"0.449"},"uses_additional_data":false,"paper_date":"2022-10-12","paper":"/paper/clustering-embedding-tables-without-first","paper_url":"https://arxiv.org/abs/2210.05974v3","paper_title":"Clustering the Sketch: A Novel Approach to Embedding Table Compression","code":"https://github.com/thomasahle/cce","n_code_links":1,"syntology":null},{"rank_in_archive_order":29,"model":"xDeepFM","metrics":{"AUC":"0.8052","Log Loss":"0.4418"},"uses_additional_data":false,"paper_date":"2018-03-14","paper":"/paper/xdeepfm-combining-explicit-and-implicit","paper_url":"http://arxiv.org/abs/1803.05170v3","paper_title":"xDeepFM: Combining Explicit and Implicit Feature Interactions for Recommender Systems","code":"https://github.com/microsoft/recommenders","n_code_links":19,"syntology":{"n_ran":3,"n_unverified":12,"n_samples":15,"n_pointer_only_licence":2}},{"rank_in_archive_order":30,"model":"WMLFF","metrics":{"AUC":"0.804","Log Loss":"0.447"},"uses_additional_data":false,"paper_date":"2023-08-03","paper":"/paper/weighted-multi-level-feature-factorization","paper_url":"https://arxiv.org/abs/2308.02568v1","paper_title":"Weighted Multi-Level Feature Factorization for App ads CTR and installation prediction","code":"https://github.com/knife982000/recsys2023challenge","n_code_links":1,"syntology":null},{"rank_in_archive_order":31,"model":"FINN","metrics":{"AUC":"0.8020","Log Loss":"0.5409"},"uses_additional_data":false,"paper_date":"2020-06-07","paper":"/paper/feature-interaction-based-neural-network-for","paper_url":"https://arxiv.org/abs/2006.05312v1","paper_title":"Feature Interaction based Neural Network for Click-Through Rate Prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":32,"model":"AutoDeepFM(3rd)","metrics":{"AUC":"0.8010","Log Loss":"0.5405"},"uses_additional_data":false,"paper_date":"2020-03-25","paper":"/paper/autofis-automatic-feature-interaction","paper_url":"https://arxiv.org/abs/2003.11235v3","paper_title":"AutoFIS: Automatic Feature Interaction Selection in Factorization Models for Click-Through Rate Prediction","code":"https://github.com/PaddlePaddle/PaddleRec/tree/master/models/rank/autofis","n_code_links":6,"syntology":{"n_ran":1,"n_unverified":23,"n_samples":24,"n_pointer_only_licence":0}},{"rank_in_archive_order":33,"model":"DeepFM","metrics":{"AUC":"0.8007","Log Loss":"0.45083"},"uses_additional_data":false,"paper_date":"2017-03-13","paper":"/paper/deepfm-a-factorization-machine-based-neural","paper_url":"http://arxiv.org/abs/1703.04247v1","paper_title":"DeepFM: A Factorization-Machine based Neural Network for CTR Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":23,"syntology":{"n_ran":2,"n_unverified":6,"n_samples":8,"n_pointer_only_licence":2}},{"rank_in_archive_order":34,"model":"TFNet","metrics":{"AUC":"0.7991"},"uses_additional_data":false,"paper_date":"2020-06-29","paper":"/paper/tfnet-multi-semantic-feature-interaction-for","paper_url":"https://arxiv.org/abs/2006.15939v1","paper_title":"TFNet: Multi-Semantic Feature Interaction for CTR Prediction","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":35,"model":"PNN*","metrics":{"AUC":"0.7987","Log Loss":"0.45214"},"uses_additional_data":false,"paper_date":"2016-11-01","paper":"/paper/product-based-neural-networks-for-user","paper_url":"http://arxiv.org/abs/1611.00144v1","paper_title":"Product-based Neural Networks for User Response Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":11,"syntology":null},{"rank_in_archive_order":36,"model":"OPNN","metrics":{"AUC":"0.7982","Log Loss":"0.45256"},"uses_additional_data":false,"paper_date":"2016-11-01","paper":"/paper/product-based-neural-networks-for-user","paper_url":"http://arxiv.org/abs/1611.00144v1","paper_title":"Product-based Neural Networks for User Response Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":11,"syntology":null},{"rank_in_archive_order":37,"model":"Wide&Deep","metrics":{"AUC":"0.7981","Log Loss":"0.46772"},"uses_additional_data":false,"paper_date":"2016-06-24","paper":"/paper/wide-deep-learning-for-recommender-systems","paper_url":"http://arxiv.org/abs/1606.07792v1","paper_title":"Wide & Deep Learning for Recommender Systems","code":"https://github.com/microsoft/recommenders","n_code_links":39,"syntology":{"n_ran":0,"n_unverified":5,"n_samples":5,"n_pointer_only_licence":5}},{"rank_in_archive_order":38,"model":"IPNN","metrics":{"AUC":"0.7972","Log Loss":"0.45323"},"uses_additional_data":false,"paper_date":"2016-11-01","paper":"/paper/product-based-neural-networks-for-user","paper_url":"http://arxiv.org/abs/1611.00144v1","paper_title":"Product-based Neural Networks for User Response Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":11,"syntology":null},{"rank_in_archive_order":39,"model":"FNN","metrics":{"AUC":"0.7963","Log Loss":"0.45738"},"uses_additional_data":false,"paper_date":"2016-01-11","paper":"/paper/deep-learning-over-multi-field-categorical","paper_url":"http://arxiv.org/abs/1601.02376v1","paper_title":"Deep Learning over Multi-field Categorical Data: A Case Study on User Response Prediction","code":"https://github.com/shenweichen/DeepCTR","n_code_links":5,"syntology":null}],"since_archive":{"claim":"Results that newer papers report for their own method, placed here by Syntology. A model pointed at the cell in the paper's own table; the number was read from that cell and checked against this leaderboard's metric, dataset, split and scale; an independent check that saw this leaderboard's other rows and every other leaderboard on the same dataset accepted it. Not reviewed by the paper's authors or by the archive's editors, and not ranked against the archive rows.","extraction_file_present":true,"measurement":{"test_papers":883,"papers_with_output":881,"judged_true":108,"judged":110,"wilson95_lower":0.9361,"measured_on":"2026-09-24","frozen_commit":"0e3de0df94"},"measurement_note":"blind adjudication of accepted entries on a held-out split of archive papers, rules frozen before the test","coverage":{"sentence":"Syntology has checked 6,264 of the 9,581 papers on this site that are newer than the archive; results from the others appear after they are checked.","complete":false,"papers_newer_than_archive":9581,"papers_checked":6264,"papers_extracted_not_yet_verified":0,"boards_without_verdict":2,"papers_not_yet_extracted":3316},"order":"newest first by month (arXiv date, else the arXiv-id month), then arXiv id descending","columns":[],"entries":[]},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":13,"rows_with_any_sample_ran":7,"distinct_papers_with_graph_line":13,"distinct_papers_with_any_sample_ran":7,"samples_over_distinct_papers":{"n_ran":24,"n_unverified":109,"n_samples":133,"n_pointer_only_licence":17,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":24,"n_unverified":109,"n_samples":133,"n_pointer_only_licence":17,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}