Papers › Survey of Big Data sizes in 2021

Survey of Big Data sizes in 2021

15 Feb 2022arXiv:2202.07659links table onlyarchive 2025-07-28

Luca Clissa

The archive published only this paper's code-link row. Authors, date and abstract are from arXiv's metadata (CC0), read from the Kaggle arXiv metadata snapshot of 2026-09-12 where its title matched the archive's; the title is the archive's.

The modern increase in data production is driven by multiple factors, and several stakeholders from various sectors contribute to it. Although drawing a comparison of the sizes at stake for different big data players is hard due to the lack of official data, this report tries to reconstruct the yearly orders of magnitude generated by some of the most important organizations by mining several online sources. The estimation is based on retrieving meaningful unitary data production measures for each of the big data sources considered, and the yearly amounts are then obtained by conjecturing reasonable per-unit sizes. The final result is summarized in the form of a bubble plot.

PaperPDFCode

Code

clissa/BigData2021 officialmentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections