Papers › Computing Web-scale Topic Models using an Asynchronous Parameter Server

Computing Web-scale Topic Models using an Asynchronous Parameter Server

24 May 2016arXiv:1605.07422archive 2025-07-28

Rolf Jagerman, Carsten Eickhoff, Maarten de Rijke

Topic models such as Latent Dirichlet Allocation (LDA) have been widely used in information retrieval for tasks ranging from smoothing and feedback methods to tools for exploratory search and discovery. However, classical methods for inferring topic models do not scale up to the massive size of today's publicly available Web-scale data sets. The state-of-the-art approaches rely on custom strategies, implementations and hardware to facilitate their asynchronous, communication-intensive workloads. We present APS-LDA, which integrates state-of-the-art topic modeling with cluster computing frameworks such as Spark using a novel asynchronous parameter server. Advantages of this integration include convenient usage of existing data processing pipelines and eliminating the need for disk writes as data can be kept in memory from start to finish. Our goal is not to outperform highly customized implementations, but to propose a general high-performance topic modeling framework that can easily be used in today's data processing pipelines. We compare APS-LDA to the existing Spark LDA implementations and show that our system can, on a 480-core cluster, process up to 135 times more data and 10 times more topics without sacrificing model quality.

PaperPDFCode

Code

rjagerman/glint officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Information RetrievalRetrievalTopic Models

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

LDA

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections