← all repositories

huggingface/optimum-quanto

A PyTorch quantization backend providing int2/int4/int8/float8 weight and activation quantization with CUDA acceleration for optimized model inference.

optimum-quanto
Not currently ranked — collecting fresh signals.
star history

Optimum Quanto is a quantization library for PyTorch models that enables dynamic and static quantization with automatic stub insertion for quantized operations and modules. It accelerates matrix multiplications on CUDA devices and supports serialization compatible with PyTorch weight_only and safetensors formats, making it useful for optimizing LLM and transformer model inference while reducing memory footprint.

Frequently asked

What is huggingface/optimum-quanto?
A PyTorch quantization backend providing int2/int4/int8/float8 weight and activation quantization with CUDA acceleration for optimized model inference.
Is optimum-quanto open source?
Yes — huggingface/optimum-quanto is open source, released under the Apache-2.0 license.
What language is optimum-quanto written in?
huggingface/optimum-quanto is primarily written in Python.
How popular is optimum-quanto?
huggingface/optimum-quanto has 1k stars on GitHub.
Where can I find optimum-quanto?
huggingface/optimum-quanto is on GitHub at https://github.com/huggingface/optimum-quanto.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.