Browse State-of-the-Art › Quantization › Papers, page 18
Quantization
Papers archive 2025-07-28
archive papers tagged: 4,925 · with a code link: 1,596 · where Syntology ran a sample: 515 (452 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (515 of 4,925 tagged: 452 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument)
Page 18 of 50: papers 1,701 to 1,800 of 4,925, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
UniHM: Universal Human Motion Generation with Object Interactions in Indoor Scenes19 May 2025 0 repositories listed
-
CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design18 May 2025 0 repositories listed
-
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies18 May 2025 0 repositories listed
-
KVmix: Gradient-Based Layer Importance-Aware Mixed-Precision Quantization for KV Cache18 May 2025 0 repositories listed
-
FedHQ: Hybrid Runtime Quantization for Federated Learning17 May 2025 0 repositories listed
-
Gaussian Weight Sampling for Scalable, Efficient and Stable Pseudo-Quantization Training16 May 2025 0 repositories listed
-
MARRS: Masked Autoregressive Unit-based Reaction Synthesis16 May 2025 0 repositories listed
-
Benchmarking CFAR and CNN-based Peak Detection Algorithms in ISAC under Hardware Impairments16 May 2025 0 repositories listed
-
Formal Uncertainty Propagation for Stochastic Dynamical Systems with Additive Noise16 May 2025 0 repositories listed
-
Qronos: Correcting the Past by Shaping the Future... in Post-Training Quantization16 May 2025 0 repositories listed
-
A probabilistic framework for dynamic quantization15 May 2025 0 repositories listed
-
VQ-Logits: Compressing the Output Bottleneck of Large Language Models via Vector Quantized Logits15 May 2025 0 repositories listed
-
Zero-shot Quantization: A Comprehensive Survey14 May 2025 0 repositories listed
-
Multi-Layer Hierarchical Federated Learning with Quantization13 May 2025 0 repositories listed
-
Resource-Efficient Language Models: Quantization for Fast and Accessible Inference13 May 2025 0 repositories listed
-
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits12 May 2025 0 repositories listed
-
Bang for the Buck: Vector Search on Cloud CPUs12 May 2025 0 repositories listed
-
Cognitive Non-Coherent Jamming Techniques for Frequency Selective Attacks12 May 2025 0 repositories listed
-
Efficient ANN-SNN Conversion with Error Compensation Learning12 May 2025 0 repositories listed
-
Private LoRA Fine-tuning of Open-Source LLMs with Homomorphic Encryption12 May 2025 0 repositories listed
-
QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads12 May 2025 0 repositories listed
-
Semantic Retention and Extreme Compression in LLMs: Can We Have Both?12 May 2025 0 repositories listed
-
Improving Block-Wise LLM Quantization by 4-bit Block-Wise Optimal Float (BOF4): Analysis and Variations10 May 2025 0 repositories listed
-
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference9 May 2025 0 repositories listed
-
LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation Quantization9 May 2025 0 repositories listed
-
Turbo-ICL: In-Context Learning-Based Turbo Equalization9 May 2025 0 repositories listed
-
Learning from Loss Landscape: Generalizable Mixed-Precision Quantization via Adaptive Sharpness-Aware Gradient Aligning8 May 2025 0 repositories listed
-
Low-bit Model Quantization for Deep Neural Networks: A Survey8 May 2025 0 repositories listed
-
Mix-QSAM: Mixed-Precision Quantization of the Segment Anything Model8 May 2025 0 repositories listed
-
ReactDance: Progressive-Granular Representation for Long-Term Coherent Reactive Dance Generation8 May 2025 0 repositories listed
-
3D Gaussian Splatting Data Compression with Mixture of Priors6 May 2025 0 repositories listed
-
Lightweight Clinical Decision Support System using QLoRA-Fine-Tuned LLMs and Retrieval-Augmented Generation6 May 2025 0 repositories listed
-
PROM: Prioritize Reduction of Multiplications Over Lower Bit-Widths for Efficient CNNs6 May 2025 0 repositories listed
-
Bielik 11B v2 Technical Report5 May 2025 0 repositories listed
-
End-to-end fully-binarized network design: from Generic Learned Thermometer to Block Pruning5 May 2025 0 repositories listed
-
EntroLLM: Entropy Encoded Weight Compression for Efficient Large Language Model Inference on Edge Devices5 May 2025 0 repositories listed
-
NeuroSim V1.5: Improved Software Backbone for Benchmarking Compute-in-Memory Accelerators with Device and Circuit-level Non-idealities5 May 2025 0 repositories listed
-
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques5 May 2025 0 repositories listed
-
Rapid yet accurate Tile-circuit and device modeling for Analog In-Memory Computing5 May 2025 0 repositories listed
-
RobSurv: Vector Quantization-Based Multi-Modal Learning for Robust Cancer Survival Prediction5 May 2025 0 repositories listed
-
Quantizing Diffusion Models from a Sampling-Aware Perspective4 May 2025 0 repositories listed
-
PASCAL: Precise and Efficient ANN- SNN Conversion using Spike Accumulation and Adaptive Layerwise Activation3 May 2025 0 repositories listed
-
Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth2 May 2025 0 repositories listed
-
Efficient Vision-based Vehicle Speed Estimation2 May 2025 0 repositories listed
-
Grouped Sequency-arranged Rotation: Optimizing Rotation Transformation for Quantization for Free2 May 2025 0 repositories listed
-
LMDepth: Lightweight Mamba-based Monocular Depth Estimation for Real-World Deployment2 May 2025 0 repositories listed
-
Aggregating empirical evidence from data strategy studies: a case on model quantization1 May 2025 0 repositories listed
-
Generative QoE Modeling: A Lightweight Approach for Telecom Networks30 Apr 2025 0 repositories listed
-
Optimization of embeddings storage for RAG systems using quantization and dimensionality reduction techniques30 Apr 2025 0 repositories listed
-
Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models30 Apr 2025 0 repositories listed
-
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech29 Apr 2025 0 repositories listed
-
Clustering-Based Evolutionary Federated Multiobjective Optimization and Learning29 Apr 2025 0 repositories listed
-
FineQ: Software-Hardware Co-Design for Low-Bit Fine-Grained Mixed-Precision Quantization of LLMs28 Apr 2025 0 repositories listed
-
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate28 Apr 2025 0 repositories listed
-
Pushing the boundary on Natural Language Inference25 Apr 2025 0 repositories listed
-
Fast Autoregressive Models for Continuous Latent Generation24 Apr 2025 0 repositories listed
-
On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration24 Apr 2025 0 repositories listed
-
Precision Neural Network Quantization via Learnable Adaptive Modules24 Apr 2025 0 repositories listed
-
Distributed Optimization with Efficient Communication, Event-Triggered Solution Enhancement, and Operation Stopping23 Apr 2025 0 repositories listed
-
Hexcute: A Tile-based Programming Language with Automatic Layout and Task-Mapping Synthesis22 Apr 2025 0 repositories listed
-
TeLLMe: An Energy-Efficient Ternary LLM Accelerator for Prefilling and Decoding on Edge FPGAs22 Apr 2025 0 repositories listed
-
Compute-Optimal LLMs Provably Generalize Better With Scale21 Apr 2025 0 repositories listed
-
StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models21 Apr 2025 0 repositories listed
-
Efficient Implicit Neural Compression of Point Clouds via Learnable Activation in Latent Space20 Apr 2025 0 repositories listed
-
FGMP: Fine-Grained Mixed-Precision Weight and Activation Quantization for Hardware-Accelerated LLM Inference19 Apr 2025 0 repositories listed
-
Lightweight Road Environment Segmentation using Vector Quantization19 Apr 2025 0 repositories listed
-
From Large to Super-Tiny: End-to-End Optimization for Cost-Efficient LLMs18 Apr 2025 0 repositories listed
-
Gradual Binary Search and Dimension Expansion : A general method for activation quantization in LLMs18 Apr 2025 0 repositories listed
-
The Binary and Ternary Quantization Can Improve Feature Discrimination18 Apr 2025 0 repositories listed
-
D²MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving17 Apr 2025 0 repositories listed
-
FedX: Adaptive Model Decomposition and Quantization for IoT Federated Learning17 Apr 2025 0 repositories listed
-
ESC-MVQ: End-to-End Semantic Communication With Multi-Codebook Vector Quantization16 Apr 2025 0 repositories listed
-
Résumé abstractif à partir d'une transcription audio16 Apr 2025 0 repositories listed
-
CSPLADE: Learned Sparse Retrieval with Causal Language Models15 Apr 2025 0 repositories listed
-
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM15 Apr 2025 0 repositories listed
-
Neural Network Emulation of the Classical Limit in Quantum Systems via Learned Observable Mappings15 Apr 2025 0 repositories listed
-
Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization13 Apr 2025 0 repositories listed
-
Simultaneous Input and State Estimation under Output Quantization: A Gaussian Mixture approach13 Apr 2025 0 repositories listed
-
Asymptotic stabilization under homomorphic encryption: A re-encryption free method12 Apr 2025 0 repositories listed
-
Deploying Large AI Models on Resource-Limited Devices with Split Federated Learning12 Apr 2025 0 repositories listed
-
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization11 Apr 2025 0 repositories listed
-
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer11 Apr 2025 0 repositories listed
-
Muon-Accelerated Attention Distillation for Real-Time Edge Synthesis via Optimized Latent Diffusion11 Apr 2025 0 repositories listed
-
SpecEE: Accelerating Large Language Model Inference with Speculative Early Exiting11 Apr 2025 0 repositories listed
-
APSQ: Additive Partial Sum Quantization with Algorithm-Hardware Co-Design10 Apr 2025 0 repositories listed
-
PoGO: A Scalable Proof of Useful Work via Quantized Gradient Descent and Merkle Proofs10 Apr 2025 0 repositories listed
-
BBQRec: Behavior-Bind Quantization for Multi-Modal Sequential Recommendation9 Apr 2025 0 repositories listed
-
CHIME: A Compressive Framework for Holistic Interest Modeling9 Apr 2025 0 repositories listed
-
AccLLM: Accelerating Long-Context LLM Inference Via Algorithm-Hardware Co-Design7 Apr 2025 0 repositories listed
-
Achieving binary weight and activation for LLMs using Post-Training Quantization7 Apr 2025 0 repositories listed
-
Balancing Robustness and Efficiency in Embedded DNNs Through Activation Function Selection7 Apr 2025 0 repositories listed
-
Bridging the Gap between Continuous and Informative Discrete Representations by Random Product Quantization7 Apr 2025 0 repositories listed
-
Two is Better than One: Efficient Ensemble Defense for Robust and Compact Models7 Apr 2025 0 repositories listed
-
Skin Color Measurement from Dermatoscopic Images: An Evaluation on a Synthetic Dataset6 Apr 2025 0 repositories listed
-
Autoregressive High-Order Finite Difference Modulo Imaging: High-Dynamic Range for Computer Vision Applications5 Apr 2025 0 repositories listed
-
Efficient FPGA-accelerated Convolutional Neural Networks for Cloud Detection on CubeSats4 Apr 2025 0 repositories listed
-
Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions4 Apr 2025 0 repositories listed
-
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency4 Apr 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.