Papers › OnPrem.LLM: A Privacy-Conscious Document Intelligence Toolkit

OnPrem.LLM: A Privacy-Conscious Document Intelligence Toolkit

12 May 2025arXiv:2505.07672archive 2025-07-28

Arun S. Maiya

We present OnPrem$.$LLM, a Python-based toolkit for applying large language models (LLMs) to sensitive, non-public data in offline or restricted environments. The system is designed for privacy-preserving use cases and provides prebuilt pipelines for document processing and storage, retrieval-augmented generation (RAG), information extraction, summarization, classification, and prompt/output processing with minimal configuration. OnPrem$.LLM supports multiple LLM backends – including llama.cpp, Ollama, vLLM, and Hugging Face Transformers – with quantized model support, GPU acceleration, and seamless backend switching. Although designed for fully local execution, OnPrem.$LLM also supports integration with a wide range of cloud LLM providers when permitted, enabling hybrid deployments that balance performance with data control. A no-code web interface extends accessibility to non-technical users.

PaperPDFCode

Code

amaiya/onprem officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Privacy PreservingRAGRetrievalRetrieval-augmented Generation

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections