{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/polisis-automated-analysis-and-presentation","title":"Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning","arxiv_id":"1802.02561","date":"2018-02-07","proceeding":null,"authors":["Hamza Harkous","Kassem Fawaz","Rémi Lebret","Florian Schaub","Kang G. Shin","Karl Aberer"],"abstract":"Privacy policies are the primary channel through which companies inform users\nabout their data collection and sharing practices. These policies are often\nlong and difficult to comprehend. Short notices based on information extracted\nfrom privacy policies have been shown to be useful but face a significant\nscalability hurdle, given the number of policies and their evolution over time.\nCompanies, users, researchers, and regulators still lack usable and scalable\ntools to cope with the breadth and depth of privacy policies. To address these\nhurdles, we propose an automated framework for privacy policy analysis\n(Polisis). It enables scalable, dynamic, and multi-dimensional queries on\nnatural language privacy policies. At the core of Polisis is a privacy-centric\nlanguage model, built with 130K privacy policies, and a novel hierarchy of\nneural-network classifiers that accounts for both high-level aspects and\nfine-grained details of privacy practices. We demonstrate Polisis' modularity\nand utility with two applications supporting structured and free-form querying.\nThe structured querying application is the automated assignment of privacy\nicons from privacy policies. With Polisis, we can achieve an accuracy of 88.4%\non this task. The second application, PriBot, is the first freeform\nquestion-answering system for privacy policies. We show that PriBot can produce\na correct answer among its top-3 results for 82% of the test questions. Using\nan MTurk user study with 700 participants, we show that at least one of\nPriBot's top-3 answers is relevant to users for 89% of the test questions.","url_abs":"http://arxiv.org/abs/1802.02561v2","url_pdf":"http://arxiv.org/pdf/1802.02561v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"polisis-automated-analysis-and-presentation","repo_url":"https://github.com/SmartDataAnalytics/Polisis_Benchmark","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"polisis-automated-analysis-and-presentation","repo_url":"https://github.com/wi-pi/GDPR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1802.02561","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}