{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/streaming-algorithms-for-diversity","title":"Streaming Algorithms for Diversity Maximization with Fairness Constraints","arxiv_id":"2208.00194","date":"2022-07-30","proceeding":null,"authors":["Yanhao Wang","Francesco Fabbri","Michael Mathioudakis"],"abstract":"Diversity maximization is a fundamental problem with wide applications in data summarization, web search, and recommender systems. Given a set $X$ of $n$ elements, it asks to select a subset $S$ of $k \\ll n$ elements with maximum \\emph{diversity}, as quantified by the dissimilarities among the elements in $S$. In this paper, we focus on the diversity maximization problem with fairness constraints in the streaming setting. Specifically, we consider the max-min diversity objective, which selects a subset $S$ that maximizes the minimum distance (dissimilarity) between any pair of distinct elements within it. Assuming that the set $X$ is partitioned into $m$ disjoint groups by some sensitive attribute, e.g., sex or race, ensuring \\emph{fairness} requires that the selected subset $S$ contains $k_i$ elements from each group $i \\in [1,m]$. A streaming algorithm should process $X$ sequentially in one pass and return a subset with maximum \\emph{diversity} while guaranteeing the fairness constraint. Although diversity maximization has been extensively studied, the only known algorithms that can work with the max-min diversity objective and fairness constraints are very inefficient for data streams. Since diversity maximization is NP-hard in general, we propose two approximation algorithms for fair diversity maximization in data streams, the first of which is $\\frac{1-\\varepsilon}{4}$-approximate and specific for $m=2$, where $\\varepsilon \\in (0,1)$, and the second of which achieves a $\\frac{1-\\varepsilon}{3m+2}$-approximation for an arbitrary $m$. Experimental results on real-world and synthetic datasets show that both algorithms provide solutions of comparable quality to the state-of-the-art algorithms while running several orders of magnitude faster in the streaming setting.","url_abs":"https://arxiv.org/abs/2208.00194v1","url_pdf":"https://arxiv.org/pdf/2208.00194v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"streaming-algorithms-for-diversity","repo_url":"https://github.com/yhwang1990/code-fdm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"data-summarization","task_name":"Data Summarization"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"fairness","task_name":"Fairness"},{"task_slug":"recommendation-systems","task_name":"Recommendation Systems"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2208.00194","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}