{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/co-design-hardware-and-algorithm-for-vector","title":"Co-design Hardware and Algorithm for Vector Search","arxiv_id":"2306.11182","date":"2023-06-19","proceeding":null,"authors":["Wenqi Jiang","Shigang Li","Yu Zhu","Johannes De Fine Licht","Zhenhao He","Runbin Shi","Cedric Renggli","Shuai Zhang","Theodoros Rekatsinas","Torsten Hoefler","Gustavo Alonso"],"abstract":"Vector search has emerged as the foundation for large-scale information retrieval and machine learning systems, with search engines like Google and Bing processing tens of thousands of queries per second on petabyte-scale document datasets by evaluating vector similarities between encoded query texts and web documents. As performance demands for vector search systems surge, accelerated hardware offers a promising solution in the post-Moore's Law era. We introduce \\textit{FANNS}, an end-to-end and scalable vector search framework on FPGAs. Given a user-provided recall requirement on a dataset and a hardware resource budget, \\textit{FANNS} automatically co-designs hardware and algorithm, subsequently generating the corresponding accelerator. The framework also supports scale-out by incorporating a hardware TCP/IP stack in the accelerator. \\textit{FANNS} attains up to 23.0$\\times$ and 37.2$\\times$ speedup compared to FPGA and CPU baselines, respectively, and demonstrates superior scalability to GPUs, achieving 5.5$\\times$ and 7.6$\\times$ speedup in median and 95\\textsuperscript{th} percentile (P95) latency within an eight-accelerator configuration. The remarkable performance of \\textit{FANNS} lays a robust groundwork for future FPGA integration in data centers and AI supercomputers.","url_abs":"https://arxiv.org/abs/2306.11182v3","url_pdf":"https://arxiv.org/pdf/2306.11182v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"co-design-hardware-and-algorithm-for-vector","repo_url":"https://github.com/WenqiJiang/SC-ANN-FPGA","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":null,"task_name":"CPU"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}