Skip to content

Set the quant_method to auto-round by default - #2138

Open
Zhenzhong1 wants to merge 6 commits into
mainfrom
zhenzhong/fix_arfp8_format
Open

Set the quant_method to auto-round by default#2138
Zhenzhong1 wants to merge 6 commits into
mainfrom
zhenzhong/fix_arfp8_format

Conversation

@Zhenzhong1

@Zhenzhong1 Zhenzhong1 commented Aug 10, 2026

Copy link
Copy Markdown
Contributor

Set the quant_method to auto-round by default.

test:
https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut

Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>
Signed-off-by: Zhenzhong1 <zhenzhong.xu@intel.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Sets FP8 exports to use auto-round as the default quantization method while preserving backend metadata.

Changes:

  • Passes FP8 format through backend.
  • Updates expected quantization metadata in CUDA tests.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

File Description
auto_round/export/formats/backends/fp8.py Preserves FP8 backend while defaulting the quantization method.
test/unit/test_cuda/export/test_auto_round_format.py Updates FP8 metadata expectations.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread auto_round/export/formats/backends/fp8.py
@AutoRoundBot

Copy link
Copy Markdown
Collaborator

/azp run Unit-Test-CUDA-AutoRound

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@yiliu30 yiliu30 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@chensuyue chensuyue added this to the 0.15.0 milestone Aug 11, 2026
@Zhenzhong1

Zhenzhong1 commented Aug 12, 2026

Copy link
Copy Markdown
Contributor Author

Ready for review and merge. @chensuyue

Test: https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut

@chensuyue

Copy link
Copy Markdown
Contributor

Ready for review and merge. @chensuyue

Test: https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut

Where did you use this model? Can you comments on the model card?

@Zhenzhong1

Copy link
Copy Markdown
Contributor Author

Ready for review and merge. @chensuyue
Test: https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut

Where did you use this model? Can you comments on the model card?

https://github.com/vllm-project/vllm/pull/47434/changes#diff-788955606955d2c9ae3030a7463a436f9e324f8e2054951a01c6e3ea222fef92R81-R88

OK. will be used in INC vLLM Path.

@chensuyue

Copy link
Copy Markdown
Contributor

Ready for review and merge. @chensuyue
Test: https://huggingface.co/INC4AI/Qwen3-8B-fp-w8g128x128-for-ut

Where did you use this model? Can you comments on the model card?

https://github.com/vllm-project/vllm/pull/47434/changes#diff-788955606955d2c9ae3030a7463a436f9e324f8e2054951a01c6e3ea222fef92R81-R88

OK. will be used in INC vLLM Path.

Please add some comments on model card, so we will keep this model.

@AutoRoundBot

Copy link
Copy Markdown
Collaborator

/azp run Unit-Test-CUDA-AutoRound

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@AutoRoundBot

Copy link
Copy Markdown
Collaborator

/azp run Unit-Test-CUDA-AutoRound

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants