/QLLM

A general 2-8 bits quantization toolbox with GPTQ/AWQ/HQQ, and export to onnx/onnx-runtime easily.

Primary LanguagePythonApache License 2.0Apache-2.0

Watchers