[Cherry-pick] PRs #2359 #2202 #2317 #2289 #2314 #2371 #2403 #2405 #2319 #2203 #2413 #2412 #2231 #2439 #2436 #2432 #2410 #2414 #2460 #2451 #2336 #2471 #2469 #2480 #2452 #2438 #2402 #2434 - #2502
Closed
chadvoegele wants to merge 28 commits into
Conversation
…thers (#2359) Type of change: Perf enhancement `promote_static_block_weight_quantizers` runs at the end of `max_calibrate`. All it does is read quantizer state -- it takes the weight the iterator hands it and throws it away. For this it goes through `iter_weights_for_calibration`, and on a fused-MoE module that iterator yields `weight[idx]`, once per expert. Under FSDP2 the fused weight is a DTensor spread across ranks, so `weight[idx]` isn't a cheap view. Slicing it makes PyTorch pull the whole fused expert weight back from every rank, just to hand over one slice that the loop then drops. That happens once per expert, per projection, per layer -- tens of thousands of round-trips on a large MoE. The fix adds `iter_weight_quantizers_for_calibration`: - The base `QuantModule` implementation delegates to `iter_weights_for_calibration` and drops the weight, so every subclass gets a correct implementation for free and the two iterators cannot drift apart. - Only the fused-MoE class — the one where materializing the weight view is itself expensive — overrides it, walking the per-expert quantizer `ModuleList` directly. It keeps the same skip condition as the weight iterator, since fetching an attribute is free and only indexing collectives. `promote_static_block_weight_quantizers` then iterates quantizers instead of `(weight, quantizer)` pairs. No other caller changes: the other four call sites genuinely use the weight. Why not just wrap the promote loop in `enable_weight_access_and_writeback`, the way `weight_only_quantize` does? It would still gather per module for weights the loop never reads. `weight_only_quantize` needs the window because it actually computes amax from the weight; promote only needs the quantizer. Also included: a warn-once check if `_amax` is ever a `DTensor`. It should not be — `_amax` is a registered buffer, and `fully_shard` shards parameters, not buffers — which is why the `reduce_amax` in this loop stays local. If that assumption ever breaks, the reduction becomes a per-quantizer collective too, and the warning says so rather than letting it degrade silently. Measured on Qwen-3.8 2.4T at world 64 (8 nodes, B300), NVFP4: | | pre-export | |---|---| | before | 79.3 min | | after | **19.0 min** | 60.3 min saved, 4.2x. The block is gone rather than shortened -- the run logs 13 silent minutes in total, all of it model load. Calibration itself is unchanged at 2:23, so the saving comes from the promote loop rather than from work moving elsewhere. End-to-end projects to 41.7 min. No API change. `mtq.quantize` picks this up automatically. - Is this change backward compatible?: ✅ Additive; `iter_weights_for_calibration` and all its callers that use the weight are untouched. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ Under 0.47 Bug Fixes — the call site dates to 0.46. - Did you get Claude approval on this PR?: ❌ Not yet. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **Performance** - Reduced calibration overhead for FSDP2-sharded fused-MoE quantization by avoiding unnecessary expert-weight gathering when only quantizer state is needed. - Preserved quantizer ordering and projection filtering during calibration. - **Bug Fixes** - Added a warning when global amax reduction requires collective processing for individual quantizers. - **Tests** - Added coverage for gated and non-gated fused-expert calibration, including verification that fused weights are not unnecessarily indexed. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: New feature Adds PTQ support for **Step-3.7** (`stepfun-ai/Step-3.7-Flash`). Follow-up to [NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665) / [OMNIML-5583](https://jirasw.nvidia.com/browse/OMNIML-5583): with the export crash fixed in #2071 the run completes, but the checkpoint it writes is silently unquantized — ```json {"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}} ``` Two independent causes, both from Step's `trust_remote_code` modeling code. **1. The expert weights were invisible to quantization.** Step-3.5 and Step-3.7 ship the same custom `MoELinear`: a plain `nn.Module` holding one 3-D `weight` of `[num_experts, out_features, in_features]`, whose `forward(x, expert_id)` runs `F.linear` against the selected slice. It is not an `nn.Linear`, and the weights sit on the projection submodule rather than on the expert container, so neither the plain-linear path nor `_fused_experts_wrapper_class` (which wants a 3-D `down_proj` *Parameter*) claims it. The `_QuantMoELinear` wrapper that handles exactly this layout has existed since #1063, but its registration was gated on the Step-3.5 class names: ```python if type(model).__name__ not in ("Step3p5ForCausalLM", "Step3p5Model"): return for module in model.modules(): if type(module).__name__ == "Step3p5MoEMLP": ``` Step-3.7's root is `Step3p7ForConditionalGeneration` and its container is `Step3p7MoEMLP`, so it returned immediately and no expert ever got a quantizer. Detection is now **structural** — a 3-D `weight` plus `num_experts` / `in_features` / `out_features` and a two-positional-argument forward — so any Step revision (or another model shipping this layout) is picked up without a third hardcoded name. `_reconstruct_fused_moe_linear` likewise matches the wrapper type instead of the generated `QuantMoELinear` class name; a model whose class is spelled differently would otherwise quantize fine but export unusable per-expert keys. **2. Step's module names don't match the general recipes.** The MoE block is `moe` and the dense sibling is `share_expert`, so `*.experts.*`, `*block_sparse_moe*` and `*mlp*` reach none of the routed experts. This PR ships `huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast` and `huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`, which select `*moe*` and disable the router (`moe.gate`) and the shared expert — mirroring the existing Step-3.5 recipe — and documents the naming trap in `modelopt_recipes/ptq.md`. ```bash python examples/hf_ptq/hf_ptq.py --model /local/Step-3.7-Flash --trust_remote_code \ --recipe huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast \ --dataset /local/cnn_dailymail --calib_size 32 --export_path /local/Step-3.7-Flash-nvfp4 ``` - `tests/unit/torch/quantization/plugins/test_moe_linear.py` — structural detection (positive plus 2-D-weight / wrong-forward negatives), registration on a Step-3.7-shaped model, per-expert quantizers with calibrated amax, and reconstruction back to the 3-D parameter. Plus two export-dispatch tests added from review: the registry resolves a `Quant_SyntheticMoELinear` to `_export_moe_linear` (this one fails against the old name-keyed registration), and the handler fills an unrouted expert's input amax. - `tests/unit/recipe/test_step3p7_recipes.py` — drives both shipped recipes over a model mirroring Step's real paths (`model.language_model.layers[i].{moe,share_expert,mlp}`): routed experts NVFP4-quantized per expert, router / shared expert / dense MLP / `lm_head` per recipe scope. Ran locally (torch 2.11, transformers 5.5.4 — the version in the bug report): the two new files (12 tests) plus `tests/unit/recipe`, `tests/unit/torch/quantization/plugins/`, `tests/unit/torch/export/test_export_weight.py` and `test_export_registry.py` — 324 passed. Full `tests/unit` (minus onnx/puzzletron): 2490 passed, with 4 pre-existing `test_quant_aware_conversion.py` failures that reproduce unchanged on clean `main`. Not run end-to-end on the real Step-3.7-Flash checkpoint (1.4 TB / 8×B200) — QA can re-run against this branch with the recipe above. - Is this change backward compatible?: ✅ — Step-3.5 keeps working; the name gate is replaced by a superset. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ - **Export dispatch keyed on the class name** (caught in review): registration was made class-name-independent, but `_export_moe_linear` in `hf_export_handlers.py` was still registered for the literal string `"QuantMoELinear"`, so a compatible class under another name bypassed the input-amax fallback for unrouted experts. The predicate now matches `_QuantMoELinear` through the MRO (lazy import, since the wrapper lives in the optional transformers plugin), keeping the name check as a fallback so the synthetic stand-ins in `test_export_registry.py` / `test_export_weight.py` still match. This was the same class-name coupling already fixed in `_reconstruct_fused_moe_linear`, in the file I hadn't looked at. - Canonical 2026 license header on the new test file. - Recipe comments corrected: `*moe*` matches the router, but **not** `share_expert` (`layers.N.share_expert.*` contains no `moe` segment) — that entry is an explicit guard, not an override. - `ptq.md` recommendation scoped to Step-3.7, since Step-3.5 has its own recipe. Pairs with #2203 (fail fast when a quant config matches no weight quantizer), which turns this class of silent no-op into an error for any model. Independent branches; either can merge first. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **New Features** - Added post-training quantization (PTQ) support for Step-3.7 Flash models, including per-expert quantization for routed MoE layers. - Added NVFP4 recipes for expert-only and MLP-only quantization, with optional FP8 KV-cache support. - **Bug Fixes** - Improved detection, quantization, reconstruction, and export of expert-indexed MoE layers across Step model revisions. - Added safeguards to prevent quantization of unsupported offloaded expert weights. - **Documentation** - Clarified Step-3.7 recipe selection, calibration guidance, and quantization exclusions. - **Tests** - Added coverage for expert, MLP, router, attention, shared-expert, export, and calibration behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix
Fix ONNX AutoCast for models whose external initializers exceed the
in-memory protobuf limit.
- Keep external tensor payloads reference-only during graph sanitization
and type inference, then materialize them once before value-dependent
classification and conversion.
- Duplicate shared initializers directly in `GraphProto`, preserving
external-data metadata without reading tensor bytes.
- Use file-backed ONNX paths for validation, shape inference, reference
execution, and custom-operator inspection when required.
- Avoid a redundant sanitizer pass in the fully sanitized AutoCast path
while preserving existing behavior for direct `PrecisionConverter` and
`convert_to_f16()` callers.
No CLI flags, dependencies, or public return types change.
```bash
python -m modelopt.onnx.autocast \
--onnx_path model.onnx \
--output_path model_bf16.onnx \
--low_precision_type bf16
```
- Ran the complete CPU-only AutoCast unit suite with no GPU visible: 248
passed.
- Ran focused ONNX utility regressions covering shared initializer
duplication and file-backed protobuf routing: 8 passed.
- Ran pre-commit on all changed files.
- Ran CPU-only integration coverage with an exact 2,147,485,696-byte
external initializer using protobuf 7.35.1 and 6.33.6.
- Verified BF16, FP16, shared-initializer, and aggregate-external-data
runtime-cast cases.
- Verified every integration output with `onnx.checker.check_model(...,
full_check=True)`; outputs expected to remain external-data-backed did
so.
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌
> 🤖 _Generated by Codex (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
- **Bug Fixes**
- Fixed ONNX AutoCast failures for models with external initializers
larger than 2 GiB.
- Improved handling of large or external-data models during shape
inference, conversion, and runtime validation.
- Preserved external initializer metadata while avoiding unnecessary
data materialization.
- Improved temporary-file cleanup when model loading or inference fails.
- Improved processing of shared initializers, nested graphs, and custom
nodes.
- **Tests**
- Added coverage for large models, external initializers, nested graphs,
custom nodes, and runtime cleanup.
- **Documentation**
- Updated the changelog with recent fixes and benchmarking information.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…LM-capable bases in merge_lora (#2289) ### What does this PR do? Type of change: New feature + bug fix Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose config nests its text dims. **1. `config_overrides` for checkpoints whose `text_config` dims don't propagate.** Some multimodal checkpoints carry the real text-tower dims only under `config.text_config`, leaving the parent fields `None`. `from_pretrained` then builds a text tower with the wrong shape. `load_vlm_or_llm` gains an optional `config_overrides` dict applied to *both* the parent config and its `text_config` before instantiation, and the three entrypoints that load checkpoints — `ar_validate.py`, `export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides` passthrough. `main.py` threads it from `ModelArguments`. **2. `merge_lora.py` could not merge into any VLM base.** It loaded via `AutoModelForCausalLM`, which cannot load architectures absent from the CausalLM Auto map — every VLM base failed. It now goes through `load_vlm_or_llm`, which routes VLMs to `AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to `AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM behavior is byte-for-byte unchanged. Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is registered with `AutoConfig` before use, and dispatches that `model_type` to its model class directly — that plugin registers only a *config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised `KeyError('cosmos3_omni')` regardless of imports. The import is wrapped in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin isn't installed. ### Usage ```bash # Checkpoint whose real dims live under config.text_config python examples/speculative_decoding/scripts/ar_validate.py \ --model_path <ckpt> --trust_remote_code \ --config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}' # Same flag on export and merge python examples/speculative_decoding/scripts/export_hf_checkpoint.py \ --model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}' python examples/speculative_decoding/scripts/merge_lora.py \ --base_model_path <base> --exported_lora_dir <out> --output_path <merged> \ --config_overrides '{"num_hidden_layers": 36}' ``` ```python model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36}) # default None ``` ### Testing Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3 LoRA run: - **Training** — the base loads with all 36 text layers and correct dims; two 4-epoch co-training runs completed (46,816 steps each). - **Export + merge** — produced `adapter_model.safetensors` and a merged base. Verified correct by per-layer weight diff: a `start_layer=18` run changed **exactly** layers 18-35, with layers 0-17 bit-identical to the base. - **AR validation** — `--config_overrides` loads the trained checkpoint; 80/80 MT-Bench samples, AR 3.42. - **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to completion. No local unit-test run: `nvidia-modelopt` isn't installed in my checkout, so `tests/conftest.py` fails to import. Relying on CI. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — `config_overrides` defaults to `None`; the `merge_lora` loader swap keeps the same class, dtype and device_map for plain LLMs. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new dependency; `transformers_cosmos3` is an optional import guarded by `contextlib.suppress`. - Did you write any new necessary tests?: ❌ — exercising these paths needs a checkpoint with a nested `text_config`, which the unit suite has no fixture for. Happy to add one if a reviewer can point me at a small suitable model. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix if you consider it changelog-worthy. - Did you get Claude approval on this PR?: ❌ — not yet run. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added JSON-based model configuration overrides across speculative decoding, training, validation, export, and LoRA workflows. - Overrides can update primary model and text configuration settings. - Expanded support for vision-language models and Cosmos3 Omni checkpoints. - **Bug Fixes** - Improved configuration handling for offline loading and checkpoint-based initialization. - Restored draft-model precision during checkpoint loading and model conversion. - Added validation for malformed, unsupported, and non-finite override values. - Standardized configuration override guidance across command-line workflows. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?
Type of change: Bug fix
Fix FP8 ONNX export for BF16 models during real-weight compression
without changing the public API or the `weights_dtype="fp32"` default.
The FP8 exporter preserves BF16 initializer bits when bridging
GraphSurgeon NumPy arrays to Torch, widens BF16 values exactly to FP32
for normalization, and leaves existing FP16/FP32 handling unchanged.
Conv scales and dequantized outputs retain the source dtype, and scales
round upward when needed so serialized values cannot cause FP8 overflow.
`weights_dtype="bf16"` is accepted as a no-op only for FP8-only models
whose floating parameters are all BF16. Registered buffers do not affect
this weight-focused decision and may preserve higher-precision regions
in the exported graph. Unsupported BF16 FP8-to-FP16 and FP32 or
mixed-parameter-to-BF16 conversions are rejected with `ValueError`
before temporary export paths are created. A narrow GraphSurgeon fix
preserves integer BF16 value-info dtypes.
### Usage
```python
onnx_bytes, metadata = get_onnx_bytes_and_metadata(
quantized_fp8_model,
(sample_input,),
weights_dtype="bf16",
onnx_opset=23,
)
```
### Testing
- Seven focused CPU regressions passed: BF16 QDQ compression and integer
dtype handling, BF16-to-BF16 and FP32-to-FP16 Conv/Linear export, and
four unsupported-conversion cases.
- QDQ utilities: 31 passed; pytest 2.25s, wall 29.88s.
- FP8 MHA exporter: 6 passed; pytest 2.05s, wall 32.71s.
- Torch deploy utilities: 51 passed; pytest 8.94s, wall 25.74s.
- Torch ONNX CPU export: 36 passed; pytest 4.85s, wall 32.38s.
- Changed-file pre-commit hooks: all passed; wall 8.07s.
- Exact-head FP8 BF16 GPU workflow at `f21d62a`: exit code 0; ONNX
checker passed; 6 FP8 initializers, 3 native `DequantizeLinear` nodes,
and 12 BF16 initializers.
- Refreshed GitHub CI at `f21d62a`: 50 passed and 1 skipped. Unit, GPU,
and regression required aggregates and Codecov passed. Two ONNX example
leaves failed because the runner could not load a cuDNN sublibrary;
their dependent example aggregate consequently failed.
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
### Additional Information
- TODO: Deliver authoritative `native`/FP32/FP16/BF16 ONNX export across
all quantized formats in follow-up pull requests.
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190
<!-- Details about the change. -->
```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt --calib_size 128 --output_dir Llama-3.1-8B-Instruct_pts
python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts
trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
--tp_size 1 \
--pp_size 1 \
--host 0.0.0.0 \
--port 8000
```
PTS and SAT tested
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
N/A
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
* Corrected the PTS model restoration path.
* **Bug Fixes**
* Model export now saves the tokenizer alongside the checkpoint.
* Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
---------
Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix:6701737 The ONNX deployment path assumed that ModelProto.ByteSize() would always return a valid size. With newer protobuf versions, querying the size of a model exceeding the protobuf serialization limit can itself raise EncodeError: Failed to serialize proto. Replaced both direct size checks with the existing is_model_too_large_for_protobuf() helper. This helper handles size-query failures conservatively and checks the protobuf size limit: Shape inference now selects the external-data/file-based path when ByteSize() fails or the model is too large. Metadata creation uses the same safe check instead of raising another serialization error. The unused TWO_GB constant was also removed. ### Usage ``` python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image ``` ### Testing the above test case pass on B100 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved ONNX model size detection during shape inference and export processing. * Ensured large models consistently use the appropriate external-data handling path. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: documentation Expose `mse_calibrate` through `model_calib.__all__` so Sphinx autosummary includes the existing MSE calibration API on the generated `model_calib` reference page. The documentation configuration honors each module's curated `__all__` surface. Although `mse_calibrate` was implemented and used by the calibration dispatcher, it was missing from that surface and was therefore filtered out during API generation. ### Usage ```python from modelopt.torch.quantization.model_calib import mse_calibrate ``` ### Testing - `pre-commit run --files modelopt/torch/quantization/model_calib.py` - `git diff --check -- modelopt/torch/quantization/model_calib.py` - Generated the recursive autosummary API tree using the repository Sphinx configuration and module template. - Verified the generated RST contains `mse_calibrate`. - Rendered the focused module page and verified its function-table link, anchor, signature, and docstring. The canonical full build was unavailable in the active environment because the configured `shibuya` theme is not installed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — the existing documentation build directly exercises this declarative autosummary contract. - Did you update Changelog?: N/A — this is a documentation-visibility repair for an existing API. - Did you get Claude approval on this PR?: N/A ### Additional Information No source implementation or calibration behavior changed. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Made the MSE calibration capability publicly available for quantization workflows. - **Chores** - Increased documentation build time limits to improve reliability for longer-running builds. - Increased multi-version test time limits to better accommodate extended test runs. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: realAsma <akuriparambi@nvidia.com> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix for #2209 Fix TEGroupedMLP per-expert weight quantizer checkpoint resharding. `TEGroupedMLP` now saves its per-expert quantizer state as singleton local shards, allowing the distributed checkpoint format to retain each expert's global identity. Restore also initializes scalar `_amax` placeholders after ModelOpt extra-state restoration so distributed checkpoint loading can populate quantizer state for experts that move between ranks. This fixes restoring quantized TEGroupedMLP checkpoints across expert-parallel and tensor-parallel topology changes. Previously there was a bug that had two parts 1. TEGroupedMLP did not mark its per-expert quantizer state as singleton_local_shards. That meant the scalar weight_quantizer.<expert>._amax state was not saved with the same globally unique expert identity as the grouped-expert weights, so DCP could not reliably redistribute it across EP layouts. 2. During restore, ModelOpt’s extra-state restoration can leave _amax absent for experts that were not local on the checkpoint’s saving rank. The subsequent distributed checkpoint load then had no destination tensor to populate. ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing - `ruff format`, `ruff check`, `mypy`, `bandit`, and repository pre-commit hooks - Focused GPU regression: ```bash python3 -m pytest tests/gpu_megatron/torch/quantization/plugins/test_megatron.py \ -k te_grouped_sharded_state_dict_reshard -v Replaced the prior metadata-only TEGroupedMLP sharded-state test with an end-to-end distributed-checkpoint save/restore regression test. The new test: - Quantizes a TEGroupedMLP with per-expert NVFP4 weight quantizers. - Assigns each local expert a distinct, deterministic `_amax` based on its global expert index. - Saves both the model distributed checkpoint and sharded ModelOpt state. - Rebuilds the model under a different TP/EP topology. - Restores ModelOpt state, loads the distributed checkpoint, and verifies each target-local expert received the expected global-expert `_amax`. The parameterized test covers: - EP=2 -> EP=1 - EP=1 -> EP=2 - TP=1 -> TP=2 - TP=2 -> TP=1 ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved checkpoint restoration for grouped quantizers by initializing missing quantization statistics with compatible shapes. - Improved restoration across supported grouped quantizer configurations, including sequential groups and parallel checkpoint layouts. - Extra module state is now finalized through supported post-load callbacks when available. - Preserved populated quantized output-layer state during checkpoint operations while removing empty placeholders. - **Tests** - Expanded checkpoint resharding coverage across tensor- and expert-parallel configurations. - Added coverage for disabled, dynamic, and other grouped quantizer scenarios. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Jenny Chen <jennifchen@nvidia.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…uantizer (#2203) ### What does this PR do? Type of change: New feature (fail-fast guard; behavior change on a previously silent path) A `quant_cfg` whose module patterns don't match the model is not an error to `set_quantizer_by_cfg` — every pattern simply matches nothing. The run then calibrates, exports, and hands back a checkpoint that is silently unquantized: ```json {"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}} ``` Nothing in the run says so. It has only ever been caught by someone reading the exported `hf_quant_config.json` afterwards — most recently on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665), after a full 8×B200 calibration), and before that on MiniMax-M3, where fused-expert detection skipped the experts and an experts-only recipe matched nothing. `mtq.quantize` now compares the config's intent against the outcome and raises **before calibration**: ``` RuntimeError: The quantization config asks for weight quantization but no weight quantizer was enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no weight quantizer: *.experts.*weight_quantizer Either the patterns do not match this architecture's module names (check the model-specific recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights were never converted to quantized modules (an unsupported custom module, e.g. a trust_remote_code MoE layout). ``` Scoped to avoid false positives: - **Only configs that ask for weight quantization** are checked (an entry with `enable` and `weight_quantizer` in its pattern), so activation-only and KV-cache-only configs are unaffected. - **Intent is read from each pattern's final entry**, since `quant_cfg` entries apply in order: a pattern that is enabled and then disabled later asks for nothing by the end. - **Configs refining an already-quantized model** (weight quantizers enabled by an earlier `mtq.quantize`) are left alone. Matching goes through `conversion._match_quantizer` — the same matcher `set_quantizer_by_cfg` used to apply the config — so "did this pattern match anything?" is answered exactly as the applying code would. A local `fnmatch` diverges on the two cases that matcher handles: `SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and fused-experts names (`..._weight_quantizers.0` normalizing to `..._weight_quantizer`). ### Usage No API change. A config that would previously have produced an unquantized checkpoint now raises: ```python mtq.quantize(model, {"quant_cfg": [ {"quantizer_name": "*", "enable": False}, {"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}}, ]}, forward_loop) # RuntimeError if the model has no `experts` modules ``` ### Testing Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one per branch of the guard: patterns matching nothing raise; an activation-only config still runs; weight patterns disabled by a later entry still run; enabled-then-retracted patterns still run; `SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer names count as matched; and the already-quantized refinement path is exercised. Each was checked to be non-vacuous by removing the corresponding branch and confirming exactly that test fails. **One existing test changed.** `tests/gpu/torch/export/test_fsdp2_export.py` parametrized over `NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case ran the FSDP2 paths against an *unquantized* model, and the new guard reported it (4 GPU failures on the first CI run, all `quant_config6`; `NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps the scoped-recipe coverage. **If reviewers would rather not change that test's meaning, the alternative is to downgrade the guard to a warning — flagging it explicitly as a decision.** I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the four MLP/experts-scoped configs raise, and the other three (`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`, `NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE models (Qwen3-MoE, gpt-oss), so no other test is affected. Ran locally after rebasing onto current `main` (torch 2.11, transformers 5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271 passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which needs `hydra`): 2676 passed, with 4 pre-existing `test_quant_aware_conversion.py` failures that reproduce unchanged on clean `main`. GPU tests were not run locally (no suitable GPU); the FSDP2 change above is reasoned from the CI failure, not re-run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ❌ — deliberately. A config that previously produced a `quant_algo: null` checkpoint now raises. Any such run was already not doing what it claimed; the three scoping rules above keep intentional non-weight quantization working. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ (Backward Breaking Changes) - Did you get Claude approval on this PR?: ❌ ### Additional Information Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes the specific model that motivated this. Independent branches; either can merge first. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Quantization now detects enabled weight-quantization patterns that do not apply to any model weights and reports a clear validation error before calibration. - Broad wildcard patterns and nested quantizers are now handled correctly. - Overlapping patterns respect the final matching setting, including later disabling rules. - Existing quantized models can be refined using the parsed configuration. - Activation-only and explicitly disabled weight-quantization configurations remain supported. - Pipeline-parallel stages without targeted weights can bypass this validation when configured to do so. - **Documentation** - Documented the process-wide override for bypassing unmatched weight-quantizer validation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?
Type of change: documentation
Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:
- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.
This also removes an inaccurate source comment without changing runtime
behavior.
### Usage
```bash
python -m modelopt.onnx.quantization \
--onnx_path=model.onnx \
--quantize_mode=int8 \
--calibration_data_path=calib.npy \
--output_path=model.quant.onnx
```
### Testing
- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording
### Before your PR is "*Ready for review*"
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A
### Additional Information
Tracking: [6701308]
> 🤖 _Generated by Codex (AI agent)._
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.
- **Tests**
- Updated quantization command examples to use the current calibration
data path option.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix FP16 entropy calibration can fail for sufficiently narrow activation ranges because NumPy may construct the histogram bin edges at FP16 precision. NumPy 2.2 and later reject the resulting collapsed bin spacing, while earlier versions can silently return invalid, non-monotonic edges. The same precision issue can recur when ONNX Runtime merges later calibration batches. Losslessly widen FP16 activation values to FP32 while calculating and merging histograms in both ONNX entropy calibration paths, then restore the source dtype at the calibration-to-quantization boundary. This keeps the histogram bins stable without changing FP16 Q/DQ scale or graph dtype semantics. FP32 inputs and public APIs are unchanged. For full-range FP16 activations, ONNX Runtime can overflow while subtracting FP16 calibration endpoints before it widens the result. Retry only a non-finite FP16 scale calculation with FP32 endpoints, then cast the finite scale back to FP16. Existing finite FP16 calculations and all non-FP16 calculations continue to use ONNX Runtime's original result. The fallback intentionally patches only the `qdq_quantizer` binding used for calibrated activation ranges. Initializer and weight quantization continue to use ONNX Runtime's existing `quant_utils` path unchanged; full-range FP16 weight scaling is outside this calibration fix. The regression tests exercise both collectors across initial collection, an equal-range merge, and an expanding-range merge. They also verify the internal FP32 histogram and external FP16 calibration-range contract, including finite saturation when restoring sanitized values. A real entropy calibration test covers full-range FP16 values and verifies finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast integration verifies the same FP16 scale-type contract through the public quantization path. N/A — no API or usage change. All tests ran with CUDA hidden. - NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules: `39 passed`. - Public INT8 entropy quantization with full-range FP16 calibration data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX Runtime session loaded. - Changed-file pre-commit hooks: passed. Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ Follow-up to #1558. > 🤖 _Generated by Codex (AI agent)._ --------- Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com> Co-authored-by: Codex <codex@openai.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix AutoQuantize can measure a group of quantized expert layers at their enclosing MLP output. That enclosing module is often a plain PyTorch container and does not carry distributed-group information, so its sensitivity score was not combined across data- or expert-parallel workers. This PR obtains the distributed groups from the quantized layers when the scoring module does not provide them. It also preserves construction order for quantized modules, scoring modules, and their registered hyperparameters so every worker accumulates scores in the same order. The temporary state needed by backward-based scoring is now managed by one shared session. The session installs and removes forward patches and invocation-specific output-gradient hooks, controls parameter gradients, and restores the active quantization recipes even when scoring raises an exception. Scoring methods remain responsible for their own score calculation. ### Usage N/A — this fixes existing AutoQuantize behavior and does not add an API or flag. ### Testing - `pre-commit run --files modelopt/torch/quantization/algorithms.py tests/unit/torch/quantization/test_autoquant.py` - `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102 passed - Added a real two-rank gradient AutoQuantize test covering MoE experts scored at an enclosing MLP. - Added regressions for deterministic hyperparameter registration, per-invocation replay for reused score modules, exact `forward`-attribute restoration, partial setup rollback, and cleanup after a scoring failure. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors). - Is this change backward compatible?: ✅ — no API or checkpoint format changes; distributed sensitivity values now include the missing reduction. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A — no new feature, deprecation, breaking change, or critical release-note item. - Did you get Claude approval on this PR?: ❌ — pending review. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved quantization scoring consistency through deterministic ordering and invocation handling. * Added more reliable distributed score aggregation, including support for mixture-of-experts models. * Improved gradient-based scoring for repeated evaluations, tuple outputs, and checkpoint-compatible workflows. * Ensured model behavior and scoring state are restored after successful or failed evaluations. * Avoided unnecessary output replay when gradients are not required. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Joshua Hill <joshua.hill@baseten.co> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?
Type of change: documentation
`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.
This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.
It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.
**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.
### Usage
No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:
```python
from modelopt.recipe import load_recipe
cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```
### Testing
Docs-only change; no code paths touched, so no new or updated tests.
- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
and the existing `CHANGELOG.rst` entry, so the documented defaults
(`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
NVFP4-input-quantizer-only scope match the code.
### Before your PR is "*Ready for review*"
- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->
### Additional Information
The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit
* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Chad's Agent Type of change: documentation. Replace legacy TensorRT checkpoint export, support matrix, and engine-build instructions with `export_hf_checkpoint` and TensorRT-LLM's PyTorch backend. Preserve the existing 0.48.0 deprecation / 0.49.0 removal notice and page URL. Update the customized-model guide and deployment skill to match. Follow the linked unified HF export guide. No API changes. - `git diff --check` passed. - `uvx pre-commit run --files docs/source/deployment/1_tensorrt_llm.rst docs/source/guides/_customized_model_quantization.rst plugins/modelopt/skills/deployment/references/trtllm.md` passed. - Full Sphinx build delegated to the Docs workflow; preview expected after deployment. - Is this change backward compatible?: ✅ Documentation only; page URL retained. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A — documentation only. - Did you update Changelog?: N/A — guidance correction; no new API deprecation. - Did you get Claude approval on this PR?: ❌ Not requested yet. Removes instructions for the TensorRT backend that current TensorRT-LLM releases no longer support. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **Documentation** - Updated TensorRT-LLM deployment guidance to use `export_hf_checkpoint` with the PyTorch backend. - Clarified that this workflow does not require TensorRT engine construction. - Updated DBRX customization instructions for exporting and deploying quantized models. - Added TensorRT-LLM version requirements and links to unified Hugging Face deployment guidance. - Removed guidance for the legacy TensorRT-LLM checkpoint exporter and outdated troubleshooting steps. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix Fail fast when AutoQuantize receives non-finite output gradients, before accumulating sensitivity scores. The error names the affected module and suggests checking the model, data, and loss; it also identifies cuDNN SDPA backward on fully masked rows as one possible cause and gives an explicit retry workaround. Unlike the earlier revision, this does not disable cuDNN or change any attention backend settings. Invalid gradients are not zeroed or ignored, and there is no automatic retry. No API or recipe changes. For the reproduced cuDNN failure, the caller can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before a fresh AutoQuantize run. - AutoQuantize unit suite: **110 passed**. Coverage includes NaN and positive/negative infinity, module diagnostics, preventing invalid score accumulation, model-state cleanup, and unchanged SDPA backend settings. - Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main `8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size 8, 512 calibration samples, and `w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`: - Default backend: the new diagnostic fired at `model.language_model.layers.39.self_attn.q_proj`; cuDNN remained enabled and no quantized model was exported. Expected-error check passed. - Explicit cuDNN-disabled fresh run: both 64-batch calibration passes, all 16 scoring batches, optimization at **5.99 effective bits**, and checkpoint export completed with exit code 0. Verified all three indexed safetensors shards and quantization configuration; the index contains 93,563 tensors. - All applicable pre-commit checks and `git diff --check` passed. Contributor guidelines and security guidance followed; commits are signed and signed off. - Is this change backward compatible?: Yes; finite-gradient behavior and backend settings are unchanged. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither added. - Did you write any new necessary tests?: Yes. - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: Yes, 0.48.0 bug fixes. - Did you get Claude approval on this PR?: No; awaiting review. This improves error reporting rather than fixing the upstream cuDNN kernel. Blackwell-specificity is not established. Checkpoint deployment/reload was not tested. A calibration-only checkpoint-resume attempt completed scoring but encountered a separate `candidate_stats` KeyError. That issue is outside this patch; the successful export validation above used a fresh run without search-state resume. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> * **Bug Fixes** * AutoQuantize now fails fast when output gradients contain non-finite values, with an actionable error identifying the affected module. * Attention backend settings are preserved and restored after successful runs and failures. * Model state is restored when sensitivity scoring encounters an error during setup or execution. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…tate dump (#2410) ### What does this PR do? Type of change: Bug fix Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel. The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the flag's own default** — so the documented invocation aborted before writing any state: ``` File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()}) ValueError: invalid literal for int() with base 10: 'eagle' ``` **Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM container, where importing `modelopt.torch` fails (the full init chain pulls in omegaconf and friends). It therefore carries `_resolve_aux_layers_standalone`, a local copy of the preset logic in `common.resolve_aux_layers`. That copy implemented the `dflash` preset and explicit id lists, but never `eagle` — while `add_aux_layers_args` defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and were unaffected; only the vLLM path forked, and nothing compared the fork against its source. This PR resolves `eagle` inline, mirroring `hf_eagle.default_eagle_aux_layer_ids`. It also fixes a second defect the bug exposes: the function already had a message naming the accepted values, but it was unreachable, because `int()` raised first. An unrecognised preset now reports what it accepts instead of surfacing the raw `int()` error — which is what made the original failure opaque. ### Usage The previously-broken documented invocation now works: ```bash cd examples/speculative_decoding python collect_hidden_states/compute_hidden_states_vllm.py \ --model Qwen/Qwen2.5-0.5B-Instruct \ --input-data ../dataset/synthetic_conversations_1k.jsonl \ --output-dir /tmp/hs_vllm \ --max-seq-len 512 --tp 1 ``` `--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8` are unchanged. ### Testing Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which pins the standalone copy to the shared implementation it mirrors: - `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately including counts small enough that the `max(0, ...)` clamps collapse ids together. - A named regression case for `nvbugs/6753684`. - `dflash` and explicit-list behaviour unchanged. - Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the actionable message. - Out-of-range ids still rejected. Divergence here is silent — the dump would write plausible-looking hidden states from the *wrong* layers, surfacing much later as a poor acceptance rate. Hence pinning to the reference rather than asserting hardcoded lists alone. All 20 assertions verified and every pre-commit hook passes (`ruff`, `mypy`, `bandit`, RST lint, license headers). One caveat worth stating plainly: **pytest could not be run locally.** `tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`, which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in this machine's torch. Each assertion was executed directly against the real module instead, but CI is the first genuine pytest run. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ — strictly widens accepted input; `dflash` and explicit lists behave identically. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ — bug fix for a defect present in a previous release. - Did you get Claude approval on this PR?: ❌ — not yet run. ### Additional Information The underlying fragility is the duplicated implementation, not this one missing branch. The function's own `TODO: drop this once common.resolve_aux_layers is decoupled from the heavy modelopt.torch import chain` is the real fix; the new test narrows the gap but does not close it. Worth tracking separately if the vLLM dump is expected to keep pace with new presets. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection. * Added support for the documented `eagle` preset alongside `dflash` and explicit layer IDs. * Improved invalid-option errors to clearly list accepted formats. * Rejects `dflash` configurations when the target model has too few layers. * Continues rejecting layer IDs outside the model’s available range. * **Documentation** * Added a v0.48.0 changelog entry for the fix. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix Fix fakequant calibration for hybrid attention/Mamba models, including NVIDIA Nemotron-3-Nano, on vLLM 0.26 and 0.28. The manual calibration scheduler path previously submitted requests with empty KV-cache block tables. Hybrid models require scheduler-compatible cache state during prefill; on current vLLM releases the empty tables caused the Mamba state to use the reserved null block and calibration activations became NaN. Request cleanup also no longer matched the vLLM 0.28 execution lifecycle, which could leave request-scoped state in the persistent batch. This PR: - Allocates non-null scratch blocks for every KV-cache group using the vLLM warmup reservation policy. - Supports both the vLLM 0.28 reservation helper and the equivalent vLLM 0.26 calculation. - Passes newly allocated blocks through `new_block_ids_to_zero` when that scheduler field is available. - Validates that the calibration batch fits in the configured cache and reports how to reduce calibration demand if it does not. - Cleans up calibration requests through a zero-token scheduler step on current vLLM, with a direct cleanup fallback for older runners. - Updates the example Dockerfile to default to vLLM 0.28.0 while retaining vLLM 0.26.0 through `VLLM_VERSION`. - Documents the validated Nemotron-3-Nano NVFP4 KV-cache workflow and clarifies that reducing `--max-num-batched-tokens` is not required. ### Usage Build the default vLLM 0.28.0 image: ```bash docker build -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.28.0 . ``` Build with vLLM 0.26.0: ```bash docker build --build-arg VLLM_VERSION=0.26.0 \ -f examples/vllm_serve/Dockerfile \ -t vllm-modelopt:v0.26.0 . ``` Calibrate and serve Nemotron-3-Nano with NVFP4 KV-cache fakequant: ```bash KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \ python examples/vllm_serve/vllm_serve_fakequant.py \ <nemotron3_nano_model_path> \ --trust-remote-code --enforce-eager -tp 8 \ --max-model-len 8192 --host 0.0.0.0 --port 8000 ``` ### Testing Validated on omniml-a0 with `NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`, tensor parallel size 8, `NVFP4_KV_CFG`, `QUANT_CALIB_SIZE=512`, and `--max-model-len 8192`. No `--max-num-batched-tokens` override was used. - vLLM 0.28.0: - All 512 calibration samples completed. - No NaNs or cache-cleanup warnings were observed. - The server started and `/health` passed. - An OpenAI-compatible completion request returned coherent generated text. - vLLM 0.26.0: - Repeated the same 512-sample TP8 calibration with the official `vllm/vllm-openai:v0.26.0` image. - No NaNs were observed. - The server started, passed `/health`, and returned coherent generated text. - Docker: - Built and verified the updated vLLM 0.28.0 image. - Focused tests: - `tests/examples/vllm_serve/test_vllm_mlflow_utils.py`: 32 passed. - Cleanup failure, missing legacy API, and legacy fallback tests: 5 passed on both vLLM 0.26.0 and 0.28.0. - Repository hooks: - Targeted pre-commit hooks for every changed Python, Markdown, and Docker file: passed. - `git diff --check`: passed. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ — added focused coverage for fail-closed cleanup, exception chaining, and the legacy cleanup fallback; the full regression was also validated end to end. - Did you update Changelog?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information The change is quantization-format agnostic. It corrects the calibration scheduler and cache lifecycle rather than special-casing `NVFP4_KV_CFG` or using an NVFP4 cast path. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Added support for configuring the vLLM version through `VLLM_VERSION`, with vLLM 0.28.0 as the default. - Added calibration and serving guidance for hybrid attention/Mamba models, including Nemotron 3 Nano with NVFP4 KV-cache fake quantization. - **Bug Fixes** - Improved calibration block handling across supported vLLM versions. - Improved calibration cleanup to preserve original errors and provide reliable fallback behavior when standard cleanup is unavailable. - **Documentation** - Documented tested versions, direct installation commands, ModelOpt setup, serving options, and guidance to avoid NaNs during batched serving. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…ve (#2460) ### What does this PR do? Type of change: Documentation `dflash.yaml` tells the reader that chat templates live in `modelopt_recipes/general/speculative_decoding/chat_templates/`. That directory does not exist, and this comment is its only mention anywhere in the tree: ``` $ git grep -n 'speculative_decoding/chat_templates' modelopt_recipes/general/speculative_decoding/dflash.yaml:19: # Templates are in modelopt_recipes/general/speculative_decoding/chat_templates/ ``` Templates actually sit beside each launcher example — `tools/launcher/examples/Qwen/Qwen3-8B/chat_template_train.jinja`, `.../MiniMax/MiniMax-M2.7-DFlash/chat_template_train.jinja`, and so on. Split out of #2201, where this two-line comment was the only reason `modelopt-recipes-codeowners` was a required reviewer on a skills-documentation PR. ### Usage No behaviour change — comment only. ### Testing None needed; the file's only change is a YAML comment. `pre-commit` passes. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated speculative decoding guidance to clarify where each model’s chat template is maintained alongside its launcher example. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ye Yu <yeyu@nvidia.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix Preserves the public ONNX graph I/O types captured at the API boundary when `PrecisionConverter` wires output casts. Type inference can change the working graph's output declaration before conversion; consulting that mutated declaration caused the required cast back to the original public type to be discarded and metadata restoration to fail. The converter now derives its I/O type map from the preserved boundary metadata and uses that map when deciding whether a cast should become a public graph output. A regression test covers an FP32 output whose working declaration is inferred as FP16, and the changelog records the corrected behavior. ```python converted = convert_to_f16(model, keep_io_types=True) ``` - Ran `pytest tests/unit/onnx/autocast/test_precisionconverter.py` (186 passed). - Ran `pytest tests/unit/onnx/autocast` (249 passed). - Ran Ruff check and format validation on the changed Python files. - Verified the original minimal end-to-end reproduction with Python 3.12 and TensorRT 10.16.1.11; conversion now completes without the output-metadata restoration error. - Verified the full CLI quantization path proceeds through the formerly failing one-Q/DQ scheme and successfully benchmarks the generated TensorRT engine. Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: N/A Tracking: [6771663] <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **Bug Fixes** - Fixed ONNX FP16 conversion to preserve public graph output types when type inference changes internal declarations. - Ensured output casts are inserted correctly when preserving input/output types is enabled. - **Tests** - Added regression coverage confirming preserved output types, correct cast insertion, and valid ONNX model generation. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Ajinkya Rasane <ajinkyaashwin@gmail.com> Co-authored-by: Ajinkya Rasane <ajinkyaashwin@gmail.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix
Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:
- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.
For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.
This PR also changes shared exporter behavior:
- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.
Other model recipe configurations remain unchanged.
```bash
python quantize.py \
--model sdxl-1.0 \
--model-dtype Half \
--trt-high-precision-dtype Half \
--format fp4 \
--block-size 16 \
--batch-size 2 \
--calib-size 128 \
--n-steps 20 \
--quantized-torch-ckpt-save-path ./sdxl-fp4 \
--onnx-dir ./onnx-sdxl-fp4
```
- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
- 302 native block-scaled NVFP4 GEMM tactics;
- 38 native FP8 Conv tactics;
- no FP4 Q/K/V projections;
- all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.
Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).
Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).
- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A
Tracking: [5565357]
<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.
- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
> 🤖 _Generated by Codex (AI agent)._
---------
Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix <!-- Details about the change. --> ### Usage ```python # Add a code snippet demonstrating how to use this ``` ### Testing <!-- Mention how have you tested your change if applicable. --> ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain why. --> - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A <!--- Mandatory --> - Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory for new features or examples. --> - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ / ❌ / N/A <!--- Very short summary of changes only for new features, backward breaking changes, deprecations, or fixes for critical bugs present in previous releases. --> - Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run `/claude review`. NVIDIA org members can self-trigger for complex changes; orthogonal to CodeRabbit. --> ### Additional Information <!-- E.g. related issue. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated the NVIDIA Nemotron offline usage example to use the BF16 source-directory pipeline path. * Pipeline configuration remains unchanged. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Jennifer Chen <jennifchen@nvidia.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Fixes #1871 The pruning documentation is split between `docs/source/guides/3_pruning.rst` and `examples/pruning/README.md`, causing confusion about which is authoritative. The examples/pruning/README.md is the comprehensive, up-to-date reference covering Minitron, Puzzletron, FastNAS, support matrix, guidelines, and distillation hyperparameters. This PR adds a note to the RST guide making clear: - The README is canonical for Minitron and Puzzletron (LLM/VLM pruning) - The guide covers FastNAS for Computer Vision models Signed-off-by: Diya <diyaismahil7@gmail.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Updated the pruning guide’s introductory content for clearer separation of general guidance and the related Minitron/Puzzletron note. - Clarified that the guide focuses on FastNAS pruning for computer vision models. - Added references to the Pruning README for Minitron and Puzzletron API examples, support information, guidelines, and distillation hyperparameters. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: didi <diyaismahil7@gmail.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…re (#2480) Type of change: Bug fix `post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional post-quantization sanity-check `full_model.generate()` unguarded, directly before `export_quantized()`. Any exception raised there aborted the whole run and discarded a completed calibration without exporting a checkpoint. Root cause (traced from [NVBug 6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 / DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with `device_map="auto"`, relying on `accelerate`'s `infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU placement. On DGX Spark's unified-memory single-GPU host, that memory probe under-reports GPU capacity, so part of even an 8B model can land on CPU — and the existing fallback shrinks the GPU budget further (`* gpu_mem_percentage`), compounding it. Calibration survives this because it never invokes the real fake-quant kernel, but the post-PTQ sanity `generate()` does, and NVFP4's dynamic-block-quantize op (`modelopt/torch/quantization/tensor_quant.py`) hard-asserts `amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes there — after ~5.8 hours of calibration, before export. This PR does not attempt to fix the underlying `device_map`/memory-probing behavior (unverified without the actual hardware/logs, which weren't reachable from this environment). Instead it makes the failure mode safe: a failure in the optional sanity check now only skips that check and warns, and export always proceeds, regardless of why `generate()` failed. No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now completes export even if the post-quantization sanity `generate()` call raises. - Added `tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`, which drives `post_quantize()` with a `full_model.generate()` that raises and asserts `export_quantized()` still runs. - Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed). - Ran `pre-commit` on the changed files (`ruff-format` reformatted line wrapping only). - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude review` --> Fixes NVBug 6752977. Linked JIRA: OMNIML-5932. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> - **Bug Fixes** - Quantized checkpoint export now continues when the optional post-quantization generation check fails. - A warning is shown when the generation check cannot complete. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…2452) Type of change: Bug fix Hybrid (e.g. Nemotron-H) checkpoints saved by the `examples/megatron_bridge` scripts could not be reloaded or exported to HuggingFace: TypeError: MLPSubmodules.__init__() missing 2 required positional arguments: 'linear_fc1' and 'linear_fc2' **Root cause.** Megatron-LM's YAML writer represents a `functools.partial` via `_partial_representer`, which passes each keyword value through `represent_data`. A dataclass instance has no representer, so it falls through to `_safe_object_representer`, which emits only `{_target_, _call_}` and drops every field. The default hybrid stack spec builds its dense-MLP and MoE layers as exactly such partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec` on the provider — which is serialized into every checkpoint's `run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written with no fields at all. Both export paths hit it: `convert.sh` via `from_auto_config`, and `export_distilled_megatron_to_hf.py` via `export_ckpt → load_megatron_model`. Every hybrid provider is affected, dense or MoE. **Fix.** `set_moe_expert_layout()` stores a named, zero-argument *factory* instead. The provider already calls a callable spec at build time (`_resolve_hybrid_stack_spec`), so model construction is unchanged — only the serialized form differs: ```yaml hybrid_stack_spec: _call_: false _target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec ``` The grouped-GEMM factory is Megatron-Bridge's own `transformer_engine_hybrid_stack_spec`, so stock tooling (`scripts/conversion/convert.sh`) resolves it without importing ModelOpt. **Known limitation — the two layouts are not symmetric.** The SequentialMLP layout has no bridge-side equivalent (the upstream TE hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt target, which `instantiate` only accepts in a process that has imported `mbridge.py` and thereby run `register_allowed_target_prefix`. A SequentialMLP hybrid checkpoint therefore converts through the ModelOpt entrypoints but not through stock `convert.sh`, where it fails on the disallowed prefix instead of on `MLPSubmodules` — no regression, but that path stays broken for this one layout. The reach is narrow: `use_moe_grouped_gemm()` returns True for any architecture with a grouped-expert export rule, NemotronH included, so SequentialMLP requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge. Both spec builders also move out of `nas/plugins/megatron.py`, which never used them, into a new `utils/plugins/megatron_layer_specs.py` beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is reached by 16 test files through `tests/_test_utils/torch/megatron/models.py`, which is bridge-free. The underlying defect is upstream in `megatron/training/config/yaml_utils.py`; this only stops ModelOpt from stepping on it, so it is worth a separate Megatron-LM issue. No API change — hybrid checkpoints saved after this fix convert with the existing commands: ```bash torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \ --student_hf_path <student_hf_model_or_path> \ --megatron_path <distill_out>/checkpoints \ --hf_export_path <hf_out> \ --export_iterations all ``` Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B Nemotron-3.5-Lightning pruned+distilled run: - **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml` (the writer used for `run_config.yaml`), reloaded via `instantiate`, resolved. `moe_grouped_gemm=True` → `TELayerNormColumnParallelLinear`/`TERowParallelLinear` + `TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a saved config cannot regress. - Applying the equivalent `run_config.yaml` fix to 32 iteration checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104 experts) with populated `MLPSubmodules` / `MoESubmodules`. - **End-to-end exports**, 6 iterations, all rc 0, each producing exactly the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards / 41.5 GiB, all weights finite, drift from the base rising monotonically with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU, `convert.sh` GPU (4×GB300, TP=4), and `export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU 123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side conversion. - `ruff check` / `ruff format --check` passed on the source files before the module move. **Not yet run:** `tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) — the GPU allocation expired. It asserts the round-trip property verified manually above, but its `HybridModelProvider(num_layers=2, hidden_size=64, num_attention_heads=4)` construction is unverified. The module move is verified only by reference grep and syntax check, so please also run one test that uses `tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run either (unavailable in the environment used). - Is this change backward compatible?: ✅ behavior; note `get_te_hybrid_stack_spec` moved module (`modelopt.torch.nas.plugins.megatron` → `modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the changelog. - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ (added, not yet executed — see Testing) - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ❌ — will run `/claude review` before marking ready. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> * **New Features** * Hybrid checkpoints now record complete layer specifications in `run_config.yaml`. * Recorded specifications support conversion to Hugging Face format. * Hybrid MoE configurations support grouped-GEMM and sequential-MLP modes. * Configuration-based reconstruction preserves the selected MoE layout. * **Compatibility** * Checkpoints from earlier releases may require manually setting the hybrid layer specification before conversion. * **Tests** * Added coverage confirming hybrid specifications survive configuration serialization and can be recreated successfully. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix `TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized for dynamic-block quantizers, while the static path immediately below it has always substituted `maxbound` for zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a recipe that applies it to an *activation* quantizer — e.g. `general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets `*mlp*input_quantizer` — feeds a raw `0.0` into `NVFP4QTensor.get_activation_scaling_factor`, whose assert aborts the entire export: ``` AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj' (type=QuantLinear): activation scaling factor 0.0 not positive. ``` Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE expert — saw only zeros, so one dead layer costs the whole run at the final export step. This factors the substitution into `_sanitize_export_amax()` and calls it from both branches. Two details beyond de-duplication: - **Branch-free, so it survives a meta `amax`.** `torch.where` + `nan_to_num` both have meta kernels; `bool()` on a meta tensor raises. The layerwise and streaming export flows carry meta `amax` — `validate_attr` short-circuits on `is_meta` for exactly that reason — so only the warning is gated on a materialized tensor. - **No longer mutates calibrated state.** The old in-place `amax[amax == 0] = ...` wrote through a view of `self._amax`; `torch.where` returns a fresh tensor, so that hazard disappears. - **Warns, with a count.** The fix turns a loud failure into a silent one, and a zero amax means calibration never activated that layer — worth surfacing rather than papering over. The message reports how many entries were substituted, since per-location dedup otherwise collapses many dead experts into one uninformative message. A healthy model emits none. Scope: only the activation path is data-dependent and reachable this way. Weight-side `_amax` uses are left alone, since a weight amax of 0 would require an all-zero weight matrix. **Knowingly left as follow-up:** `export/quant_utils.py::get_scaling_factor` discards the sanitized `amax` when `num_bits == (2, 1)` and recomputes via `get_weights_scaling_factor_2_from_quantizer`, which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input* quantizer on a module whose *weight* quantizer is a different format (or disabled) can still trip `assert torch.all(scaling_factor > 0)`. Format dispatch is weight-driven, so the reported recipe does not reach that branch; fixing it properly changes a signature shared with the weight-side callers and is out of scope here. Not a regression. The dynamic early return, the `type: dynamic` numerics unit, and the recipe that combines them all ship in released 0.46.0 / 0.46.1. No new or changed API. Exports that previously aborted now complete and warn: ```python mtq.quantize(model, quant_cfg, forward_loop) export_hf_checkpoint(model, export_dir=out) # before: AssertionError; now: exports + UserWarning ``` - New `test_amax_export_unusable_amax`, parametrized over zero and NaN, covering the dynamic-NVFP4 and static per-tensor configs; asserts the exported scale is positive and that export leaves the calibrated `amax` untouched. Plus `test_amax_export_meta_amax`, pinning that a meta `amax` survives export rather than raising. Both run on CPU and CUDA via the shared tester. - `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40 passed. `tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed (GB300). - End-to-end repro on GB300, small Llama with one MLP fed all-zero activations under `general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()` `0.0` → `6.0`, live layer unchanged at `3.921875`, and `export_hf_checkpoint` goes from the `AssertionError` above to writing `model.safetensors`. - Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on a healthy model (Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal path is unaffected. - `pre-commit run` clean on all changed files. - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: ✅ - Did you get Claude approval on this PR?: ✅ — `/claude review` run; its one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs addressed or answered in 251f2e3d Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter also notes it passed on 0.47.0rc0; that is not explained by code — `git diff 0.47.0rc0..0.47.0rc1` touches `export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new `clamp_fp8_scales` argument whose default preserves the old behaviour) and the INT4-AWQ packing path, neither of which is on the dense-HF NVFP4 activation-scale path. Whether `amax` lands on exactly 0 is calibration/model-state dependent, which is what makes it look version-flaky. Worth flagging separately: in the reported log the **pre-PTQ** sample output is already gibberish, so that BF16 checkpoint looks broken independently of quantization. This change stops the crash, but such a run will now export a valid-but-garbage checkpoint — the new warning is the signal to investigate. Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing release. 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> * **Bug Fixes** * Fixed Hugging Face checkpoint export when dynamic-block quantizers have zero or invalid calibration scales. * Exports now use a positive fallback scale and issue a warning instead of failing when applicable. * Export operations no longer modify the original calibrated quantizer state. * Meta-device exports remain non-erroring and preserve device placement. * **Tests** * Added coverage for zero- and invalid-scale exports across dynamic and static quantization modes. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Yue <yueshen@nvidia.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do? Type of change: Bug fix: 6701777 Regression source: "[OMNIML-3349] Add FP8 MHA quantization support for HuggingFace ViT" (#1289), merged into 0.44.0rc3 via the batch cherry-pick #1350. This PR: 1. Registers nn.LayerNorm as a QuantModule for the first time (modelopt/torch/quantization/nn/modules/quant_layernorm.py), intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output quantizer rules apply to ViT. 2. Removes the prior forced Cast-alignment logic in export_onnx.py that used to normalize Q/DQ node dtypes to trt_high_precision_dtype. Root Cause: Once nn.LayerNorm became a registered QuantModule, these wildcards started unintentionally matching norm1.norm inside FLUX's AdaLayerNormZero block — an elementwise_affine=False LayerNorm with no learnable weight/bias. Its input got routed through NVFP4 Q/DQ (emitted as Float32) while its synthesized affine scale remained native BFloat16, producing the dtype mismatch. Chosen fix: Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets rather than touching the global QuantModuleRegistry (which ViT FP8 MHA still needs). Add, in both modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules (list order matters — later entries override earlier ones): - parent_class: 'nn.LayerNorm' quantizer_name: '*' enable: false ### Usage ``` python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16 trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 ``` ### Testing The above test commands. ### Before your PR is "*Ready for review*" Make sure you read and follow [Contributor guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md) and your commits are signed (`git commit -s -S`). Make sure you read and follow the [Security Best Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors) (e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(..., weights_only=False)`, `pickle`, etc.). - Is this change backward compatible?: N/A - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: N/A - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information N/A <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by excluding LayerNorm modules from quantization. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
… disabled quantization during calibration (#2434) ### What does this PR do? Type of change: Bug fix During max calibration, `enable_stats_collection()` calls `disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ attention paths bypass `TensorQuantizer.forward()` and previously selected the Triton/Kitchen path from the configured enabled state alone, so P quant-dequant could still execute while quantization was inactive. This change: - enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is true; - bypasses all fused P-QDQ paths when the P quantizer is disabled or quantization is inactive; - adds focused parameterized regression tests covering Triton and Kitchen dispatch across enabled, `disable()`, and `disable_quant()` states. ### Usage N/A. This restores the existing `disable_quant()` contract and does not introduce a new API. ### Testing - `pytest tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14 passed - Targeted pre-commit checks passed - Kitchen coverage verifies the full `{disable(), disable_quant()} × {lazy, already initialized}` dispatch matrix remains bypassed while inactive and resumes after re-enabling - Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides on B300: [module regression build #121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/) - Controlled B300 isolation passed when only the P-BMM quantizers were hard-disabled during calibration, and also passed when the existing single Triton attention configuration was forced - Fixed-code B300 validation passed in three independent runs: [Jenkins build 123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/), [Jenkins build 124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/), and [Jenkins build 125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/). Jenkins build 123 compared all 178 module outputs byte-identically and found 0/377 `amax` and 0/377 `scale` changes. The confirmed impact is incorrect calibration behavior plus unstable quantizer state and module outputs; downstream benchmark accuracy impact has not been established. ### Before your PR is "*Ready for review*" - Is this change backward compatible?: ✅ - If you copied code from any other sources or added a new PIP dependency, did you follow guidance in `CONTRIBUTING.md`: N/A - Did you write any new necessary tests?: ✅ - Did you update [Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?: N/A - Did you get Claude approval on this PR?: N/A ### Additional Information The failure was isolated to fused P-QDQ runtime-state dispatch. Quantizer topology and configuration were identical in the failing same-wheel comparison. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Corrected attention dispatch so disabled or inactive quantization uses the original attention implementation. - Preserved optimized quantized attention when quantization is enabled. - Ensured attention masks remain unchanged when using the original attention implementation. - Improved fallback behavior for fused attention paths, including correct initialization and reuse when quantization is re-enabled. - **Tests** - Added coverage for enabled and disabled quantization states, attention-mask handling, fallback selection, and result consistency. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com> Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
chadvoegele
requested review from
ajrasane,
h-guo18,
kevalmorabia97 and
meenchen
and removed request for
a team
September 22, 2026 17:49
Contributor
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Currently processing new changes in this PR. This may take a few minutes, please wait... ⚙️ Run configurationConfiguration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (91)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Contributor
|
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## release/0.47.0 #2502 +/- ##
==================================================
- Coverage 78.68% 78.64% -0.05%
==================================================
Files 526 527 +1
Lines 61331 61635 +304
==================================================
+ Hits 48259 48473 +214
- Misses 13072 13162 +90
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Cherry-picked PRs
Summary by CodeRabbit
New Features
Bug Fixes
Documentation