Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.
-
Updated
Mar 15, 2025 - Python
Efficient AI Backbones including GhostNet, TNT and MLP, developed by Huawei Noah's Ark Lab.
A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
EfficientFormerV2 [ICCV 2023] & EfficientFormer [NeurIPs 2022]
[CVPR 2024] DeepCache: Accelerating Diffusion Models for Free
Code for paper " AdderNet: Do We Really Need Multiplications in Deep Learning?"
List of papers related to neural network quantization in recent AI conferences and journals.
[NeurIPS 2024 Spotlight]"LightGaussian: Unbounded 3D Gaussian Compression with 15x Reduction and 200+ FPS", Zhiwen Fan, Kevin Wang, Kairun Wen, Zehao Zhu, Dejia Xu, Zhangyang Wang
[ICML 2024] SqueezeLLM: Dense-and-Sparse Quantization
Learning Efficient Convolutional Networks through Network Slimming, In ICCV 2017.
Fast, lossless LLM inference via dual-view diffusion decoding.
[NeurIPS 2024] KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
📚 Collection of awesome generation acceleration resources.
On-device LLM Inference Powered by X-Bit Quantization
Explorations into some recent techniques surrounding speculative decoding
[ECCV2022] Efficient Long-Range Attention Network for Image Super-resolution
[CVPR 2021] Exploring Sparsity in Image Super-Resolution for Efficient Inference
(CVPR 2021, Oral) Dynamic Slimmable Network
[NeurIPS 2024] AsyncDiff: Parallelizing Diffusion Models by Asynchronous Denoising
A curated list of papers and code on efficient diffusion models for image, video, world modeling, and language generation. Covering acceleration, quantization, compression, caching, and distillation.
To associate your repository with the efficient-inference topic, visit your repo's landing page and select "manage topics."