nvidia/model-optimizer is a library of model optimization techniques that compress deep learning models and export optimized checkpoints for deployment in downstream inference frameworks. It provides quantization, pruning, distillation, and other methods to accelerate inference speed.
Project overview
Bringing multiple optimization techniques into one library with integration into the NVIDIA deployment ecosystem via TensorRT and TensorRT-LLM, it addresses the practical need to compress and accelerate models before inference.
Project type
Model Development · Model Runtime · Infrastructure
Deployment
Refer to project documentation
License
Apache-2.0
Best for
AI engineers and developers who need to compress and optimize deep learning models for deployment in inference frameworks.
Teams already working within the NVIDIA AI software ecosystem who need optimized checkpoints for TensorRT or TensorRT-LLM.
Key capabilities
Compress model size by 2x-4x while preserving model quality.
Refine accuracy of quantized models even further with a few training steps.
Reduce model parameters or memory footprint and accelerate inference by removing unnecessary weights.
Reduce deployment model size by teaching small models to behave like larger models.
Train draft modules to predict extra tokens during inference.
Efficiently compress model by storing only its non-zero parameter values and their locations.
Limitations and risks
The project is in a pre-1.0 state that allows breaking changes in minor version updates with only a 1-release migration period.
Getting started
Install the package using pip install -U nvidia-modelopt[all]. Setup requires a Python environment or NVIDIA container images.
Evidence and sources
GitHub project description: A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learnin…
README: NVIDIA Model Optimizer (referred to as Model Optimizer, or ModelOpt) is a library comprising state-of-the-art model optimization techniques including quantization, pruning, Neural…
README: Model Optimizer is now open source!
README: Since Model Optimizer is still pre-1.0, we provide a 1-release (~1-month) migration period after deprecation.
README: Post Training Quantization | Compress model size by 2x-4x, speeding up inference while preserving model quality!