{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/automated-website-fingerprinting-through-deep","title":"Automated Website Fingerprinting through Deep Learning","arxiv_id":"1708.06376","date":"2017-08-21","proceeding":null,"authors":["Vera Rimmer","Davy Preuveneers","Marc Juarez","Tom Van Goethem","Wouter Joosen"],"abstract":"Several studies have shown that the network traffic that is generated by a\nvisit to a website over Tor reveals information specific to the website through\nthe timing and sizes of network packets. By capturing traffic traces between\nusers and their Tor entry guard, a network eavesdropper can leverage this\nmeta-data to reveal which website Tor users are visiting. The success of such\nattacks heavily depends on the particular set of traffic features that are used\nto construct the fingerprint. Typically, these features are manually engineered\nand, as such, any change introduced to the Tor network can render these\ncarefully constructed features ineffective. In this paper, we show that an\nadversary can automate the feature engineering process, and thus automatically\ndeanonymize Tor traffic by applying our novel method based on deep learning. We\ncollect a dataset comprised of more than three million network traces, which is\nthe largest dataset of web traffic ever used for website fingerprinting, and\nfind that the performance achieved by our deep learning approaches is\ncomparable to known methods which include various research efforts spanning\nover multiple years. The obtained success rate exceeds 96% for a closed world\nof 100 websites and 94% for our biggest closed world of 900 classes. In our\nopen world evaluation, the most performant deep learning model is 2% more\naccurate than the state-of-the-art attack. Furthermore, we show that the\nimplicit features automatically learned by our approach are far more resilient\nto dynamic changes of web content over time. We conclude that the ability to\nautomatically construct the most relevant traffic features and perform accurate\ntraffic recognition makes our deep learning based approach an efficient,\nflexible and robust technique for website fingerprinting.","url_abs":"http://arxiv.org/abs/1708.06376v2","url_pdf":"http://arxiv.org/pdf/1708.06376v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"automated-website-fingerprinting-through-deep","repo_url":"https://github.com/YasodGinige/TrafficGPT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"feature-engineering","task_name":"Feature Engineering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}