Set the quant_method to auto-round by default - #2138
Conversation
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
There was a problem hiding this comment.
Pull request overview
Sets FP8 exports to use auto-round as the default quantization method while preserving backend metadata.
Changes:
- Passes FP8 format through
backend. - Updates expected quantization metadata in CUDA tests.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
auto_round/export/formats/backends/fp8.py |
Preserves FP8 backend while defaulting the quantization method. |
test/unit/test_cuda/export/test_auto_round_format.py |
Updates FP8 metadata expectations. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
f805df5 to
a82897d
Compare
|
Ready for review and merge. @chensuyue Test: https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut |
Where did you use this model? Can you comments on the model card? |
OK. will be used in INC vLLM Path. |
Please add some comments on model card, so we will keep this model. |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
|
/azp run Unit-Test-CUDA-AutoRound |
|
Azure Pipelines successfully started running 1 pipeline(s). |
Set the quant_method to auto-round by default.
test:
https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut