Skip to content

[Cherry-pick] PRs #2359 #2202 #2317 #2289 #2314 #2371 #2403 #2405 #2319 #2203 #2413 #2412 #2231 #2439 #2436 #2432 #2410 #2414 #2460 #2451 #2336 #2471 #2469 #2480 #2452 #2438 #2402 #2434 - #2502

Closed
chadvoegele wants to merge 28 commits into
release/0.47.0from
cherry-picks/release-0.47.0
Closed

chadvoegele wants to merge 28 commits into
release/0.47.0from
cherry-picks/release-0.47.0

Conversation

@chadvoegele

@chadvoegele chadvoegele commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Cherry-picked PRs

Summary by CodeRabbit

  • New Features

    • Added Step-3.7 post-training quantization recipes for expert-only and MLP-focused workflows.
    • Added Diffusers SDXL NVFP4/FP8 convolution support and expanded FP4 export guidance.
    • Added configuration overrides for speculative decoding model loading.
    • Added support for the Eagle speculative-decoding preset.
    • Added vLLM 0.28 support and improved KV-cache calibration for hybrid models.
  • Bug Fixes

    • Improved ONNX handling for large models, BF16/FP8 exports, FP16 calibration, and external data.
    • Quantization now reports configurations that match no weight modules.
    • Post-quantization checks continue exporting when optional generation validation fails.
  • Documentation

    • Updated deployment guidance to use unified Hugging Face checkpoints with TensorRT-LLM’s PyTorch backend.

sugunav14 and others added 28 commits September 22, 2026 16:44
…thers (#2359)

Type of change: Perf enhancement

`promote_static_block_weight_quantizers` runs at the end of
`max_calibrate`. All it does is read
quantizer state -- it takes the weight the iterator hands it and throws
it away. For this it goes through `iter_weights_for_calibration`, and on
a fused-MoE module that iterator yields `weight[idx]`, once per expert.

Under FSDP2 the fused weight is a DTensor spread across ranks, so
`weight[idx]` isn't a cheap view.
Slicing it makes PyTorch pull the whole fused expert weight back from
every rank, just to hand over
one slice that the loop then drops. That happens once per expert, per
projection, per layer -- tens
of thousands of round-trips on a large MoE.

The fix adds `iter_weight_quantizers_for_calibration`:

- The base `QuantModule` implementation delegates to
`iter_weights_for_calibration` and drops
the weight, so every subclass gets a correct implementation for free and
the two iterators
  cannot drift apart.
- Only the fused-MoE class — the one where materializing the weight view
is itself expensive —
overrides it, walking the per-expert quantizer `ModuleList` directly. It
keeps the same skip
condition as the weight iterator, since fetching an attribute is free
and only indexing
  collectives.

`promote_static_block_weight_quantizers` then iterates quantizers
instead of `(weight, quantizer)`
pairs. No other caller changes: the other four call sites genuinely use
the weight.

Why not just wrap the promote loop in
`enable_weight_access_and_writeback`, the way
`weight_only_quantize` does? It would still gather per module for
weights the
loop never reads. `weight_only_quantize` needs the window because it
actually computes amax from the
weight; promote only needs the quantizer.

Also included: a warn-once check if `_amax` is ever a `DTensor`. It
should not be — `_amax` is a
registered buffer, and `fully_shard` shards parameters, not buffers —
which is why the
`reduce_amax` in this loop stays local. If that assumption ever breaks,
the reduction becomes a
per-quantizer collective too, and the warning says so rather than
letting it degrade silently.

Measured on Qwen-3.8 2.4T at world 64 (8 nodes, B300), NVFP4:

| | pre-export |
|---|---|
| before | 79.3 min |
| after | **19.0 min** |

60.3 min saved, 4.2x. The block is gone rather than shortened -- the run
logs 13 silent minutes
in total, all of it model load. Calibration itself is unchanged at 2:23,
so the saving comes from
the promote loop rather than from work moving elsewhere. End-to-end
projects to 41.7 min.

No API change. `mtq.quantize` picks this up automatically.

- Is this change backward compatible?: ✅ Additive;
`iter_weights_for_calibration` and all its
  callers that use the weight are untouched.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ Under 0.47 Bug Fixes — the call site dates to 0.46.
- Did you get Claude approval on this PR?: ❌ Not yet.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **Performance**
- Reduced calibration overhead for FSDP2-sharded fused-MoE quantization
by avoiding unnecessary expert-weight gathering when only quantizer
state is needed.
- Preserved quantizer ordering and projection filtering during
calibration.

- **Bug Fixes**
- Added a warning when global amax reduction requires collective
processing for individual quantizers.

- **Tests**
- Added coverage for gated and non-gated fused-expert calibration,
including verification that fused weights are not unnecessarily indexed.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Suguna Velury <178320438+sugunav14@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: New feature

Adds PTQ support for **Step-3.7** (`stepfun-ai/Step-3.7-Flash`).
Follow-up to [NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665) /
[OMNIML-5583](https://jirasw.nvidia.com/browse/OMNIML-5583): with the
export crash fixed in #2071 the run completes, but the checkpoint it
writes is silently unquantized —

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Two independent causes, both from Step's `trust_remote_code` modeling
code.

**1. The expert weights were invisible to quantization.** Step-3.5 and
Step-3.7 ship the same custom `MoELinear`: a plain `nn.Module` holding
one 3-D `weight` of `[num_experts, out_features, in_features]`, whose
`forward(x, expert_id)` runs `F.linear` against the selected slice. It
is not an `nn.Linear`, and the weights sit on the projection submodule
rather than on the expert container, so neither the plain-linear path
nor `_fused_experts_wrapper_class` (which wants a 3-D `down_proj`
*Parameter*) claims it.

The `_QuantMoELinear` wrapper that handles exactly this layout has
existed since #1063, but its registration was gated on the Step-3.5
class names:

```python
if type(model).__name__ not in ("Step3p5ForCausalLM", "Step3p5Model"):
    return
for module in model.modules():
    if type(module).__name__ == "Step3p5MoEMLP":
```

Step-3.7's root is `Step3p7ForConditionalGeneration` and its container
is `Step3p7MoEMLP`, so it returned immediately and no expert ever got a
quantizer. Detection is now **structural** — a 3-D `weight` plus
`num_experts` / `in_features` / `out_features` and a
two-positional-argument forward — so any Step revision (or another model
shipping this layout) is picked up without a third hardcoded name.
`_reconstruct_fused_moe_linear` likewise matches the wrapper type
instead of the generated `QuantMoELinear` class name; a model whose
class is spelled differently would otherwise quantize fine but export
unusable per-expert keys.

**2. Step's module names don't match the general recipes.** The MoE
block is `moe` and the dense sibling is `share_expert`, so
`*.experts.*`, `*block_sparse_moe*` and `*mlp*` reach none of the routed
experts. This PR ships
`huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast` and
`huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`, which select `*moe*`
and disable the router (`moe.gate`) and the shared expert — mirroring
the existing Step-3.5 recipe — and documents the naming trap in
`modelopt_recipes/ptq.md`.

```bash
python examples/hf_ptq/hf_ptq.py --model /local/Step-3.7-Flash --trust_remote_code \
    --recipe huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast \
    --dataset /local/cnn_dailymail --calib_size 32 --export_path /local/Step-3.7-Flash-nvfp4
```

- `tests/unit/torch/quantization/plugins/test_moe_linear.py` —
structural detection (positive plus 2-D-weight / wrong-forward
negatives), registration on a Step-3.7-shaped model, per-expert
quantizers with calibrated amax, and reconstruction back to the 3-D
parameter. Plus two export-dispatch tests added from review: the
registry resolves a `Quant_SyntheticMoELinear` to `_export_moe_linear`
(this one fails against the old name-keyed registration), and the
handler fills an unrouted expert's input amax.
- `tests/unit/recipe/test_step3p7_recipes.py` — drives both shipped
recipes over a model mirroring Step's real paths
(`model.language_model.layers[i].{moe,share_expert,mlp}`): routed
experts NVFP4-quantized per expert, router / shared expert / dense MLP /
`lm_head` per recipe scope.

Ran locally (torch 2.11, transformers 5.5.4 — the version in the bug
report): the two new files (12 tests) plus `tests/unit/recipe`,
`tests/unit/torch/quantization/plugins/`,
`tests/unit/torch/export/test_export_weight.py` and
`test_export_registry.py` — 324 passed. Full `tests/unit` (minus
onnx/puzzletron): 2490 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`.

Not run end-to-end on the real Step-3.7-Flash checkpoint (1.4 TB /
8×B200) — QA can re-run against this branch with the recipe above.

- Is this change backward compatible?: ✅ — Step-3.5 keeps working; the
name gate is replaced by a superset.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

- **Export dispatch keyed on the class name** (caught in review):
registration was made class-name-independent, but `_export_moe_linear`
in `hf_export_handlers.py` was still registered for the literal string
`"QuantMoELinear"`, so a compatible class under another name bypassed
the input-amax fallback for unrouted experts. The predicate now matches
`_QuantMoELinear` through the MRO (lazy import, since the wrapper lives
in the optional transformers plugin), keeping the name check as a
fallback so the synthetic stand-ins in `test_export_registry.py` /
`test_export_weight.py` still match. This was the same class-name
coupling already fixed in `_reconstruct_fused_moe_linear`, in the file I
hadn't looked at.
- Canonical 2026 license header on the new test file.
- Recipe comments corrected: `*moe*` matches the router, but **not**
`share_expert` (`layers.N.share_expert.*` contains no `moe` segment) —
that entry is an explicit guard, not an override.
- `ptq.md` recommendation scoped to Step-3.7, since Step-3.5 has its own
recipe.

Pairs with #2203 (fail fast when a quant config matches no weight
quantizer), which turns this class of silent no-op into an error for any
model. Independent branches; either can merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **New Features**
- Added post-training quantization (PTQ) support for Step-3.7 Flash
models, including per-expert quantization for routed MoE layers.
- Added NVFP4 recipes for expert-only and MLP-only quantization, with
optional FP8 KV-cache support.

- **Bug Fixes**
- Improved detection, quantization, reconstruction, and export of
expert-indexed MoE layers across Step model revisions.
- Added safeguards to prevent quantization of unsupported offloaded
expert weights.

- **Documentation**
- Clarified Step-3.7 recipe selection, calibration guidance, and
quantization exclusions.

- **Tests**
- Added coverage for expert, MLP, router, attention, shared-expert,
export, and calibration behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

Fix ONNX AutoCast for models whose external initializers exceed the
in-memory protobuf limit.

- Keep external tensor payloads reference-only during graph sanitization
and type inference, then materialize them once before value-dependent
classification and conversion.
- Duplicate shared initializers directly in `GraphProto`, preserving
external-data metadata without reading tensor bytes.
- Use file-backed ONNX paths for validation, shape inference, reference
execution, and custom-operator inspection when required.
- Avoid a redundant sanitizer pass in the fully sanitized AutoCast path
while preserving existing behavior for direct `PrecisionConverter` and
`convert_to_f16()` callers.

No CLI flags, dependencies, or public return types change.

```bash
python -m modelopt.onnx.autocast \
    --onnx_path model.onnx \
    --output_path model_bf16.onnx \
    --low_precision_type bf16
```

- Ran the complete CPU-only AutoCast unit suite with no GPU visible: 248
passed.
- Ran focused ONNX utility regressions covering shared initializer
duplication and file-backed protobuf routing: 8 passed.
- Ran pre-commit on all changed files.
- Ran CPU-only integration coverage with an exact 2,147,485,696-byte
external initializer using protobuf 7.35.1 and 6.33.6.
- Verified BF16, FP16, shared-initializer, and aggregate-external-data
runtime-cast cases.
- Verified every integration output with `onnx.checker.check_model(...,
full_check=True)`; outputs expected to remain external-data-backed did
so.

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

> 🤖 _Generated by Codex (AI agent)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **Bug Fixes**
- Fixed ONNX AutoCast failures for models with external initializers
larger than 2 GiB.
- Improved handling of large or external-data models during shape
inference, conversion, and runtime validation.
- Preserved external initializer metadata while avoiding unnecessary
data materialization.
- Improved temporary-file cleanup when model loading or inference fails.
- Improved processing of shared initializers, nested graphs, and custom
nodes.
- **Tests**
- Added coverage for large models, external initializers, nested graphs,
custom nodes, and runtime cleanup.
- **Documentation**
- Updated the changelog with recent fixes and benchmarking information.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…LM-capable bases in merge_lora (#2289)

### What does this PR do?

Type of change: New feature + bug fix

Two related gaps, both hit while enabling EAGLE3 on a checkpoint whose
config nests its text dims.

**1. `config_overrides` for checkpoints whose `text_config` dims don't
propagate.**
Some multimodal checkpoints carry the real text-tower dims only under
`config.text_config`, leaving the parent fields `None`.
`from_pretrained` then builds a text tower with the wrong shape.
`load_vlm_or_llm` gains an optional `config_overrides` dict applied to
*both* the parent config and its `text_config` before instantiation, and
the three entrypoints that load checkpoints — `ar_validate.py`,
`export_hf_checkpoint.py`, `merge_lora.py` — get a `--config_overrides`
passthrough. `main.py` threads it from `ModelArguments`.

**2. `merge_lora.py` could not merge into any VLM base.**
It loaded via `AutoModelForCausalLM`, which cannot load architectures
absent from the CausalLM Auto map — every VLM base failed. It now goes
through `load_vlm_or_llm`, which routes VLMs to
`AutoModelForVision2Seq`/`AutoModelForImageTextToText` and plain LLMs to
`AutoModelForCausalLM` with the same `dtype`/`device_map`, so LLM
behavior is byte-for-byte unchanged.

Also adds an optional `transformers_cosmos3` import so `cosmos3_omni` is
registered with `AutoConfig` before use, and dispatches that
`model_type` to its model class directly — that plugin registers only a
*config*, never a model under `Auto*`, so `AutoModelForCausalLM` raised
`KeyError('cosmos3_omni')` regardless of imports. The import is wrapped
in `contextlib.suppress(ImportError)`, so it is a no-op when the plugin
isn't installed.

### Usage

```bash
# Checkpoint whose real dims live under config.text_config
python examples/speculative_decoding/scripts/ar_validate.py \
    --model_path <ckpt> --trust_remote_code \
    --config_overrides '{"num_hidden_layers": 36, "intermediate_size": 12288, "num_key_value_heads": 8}'

# Same flag on export and merge
python examples/speculative_decoding/scripts/export_hf_checkpoint.py \
    --model_path <ckpt> --export_path <out> --config_overrides '{"num_hidden_layers": 36}'
python examples/speculative_decoding/scripts/merge_lora.py \
    --base_model_path <base> --exported_lora_dir <out> --output_path <merged> \
    --config_overrides '{"num_hidden_layers": 36}'
```

```python
model = load_vlm_or_llm(path, config_overrides={"num_hidden_layers": 36})  # default None
```

### Testing

Exercised end-to-end on a Cosmos3-Nano (16B, 36-layer text tower) EAGLE3
LoRA run:

- **Training** — the base loads with all 36 text layers and correct
dims; two 4-epoch co-training runs completed (46,816 steps each).
- **Export + merge** — produced `adapter_model.safetensors` and a merged
base. Verified correct by per-layer weight diff: a `start_layer=18` run
changed **exactly** layers 18-35, with layers 0-17 bit-identical to the
base.
- **AR validation** — `--config_overrides` loads the trained checkpoint;
80/80 MT-Bench samples, AR 3.42.
- **Regression check** — `merge_lora` via `load_vlm_or_llm` produces a
base loadable by `lm_eval`; ifeval/arc_challenge/winogrande all ran to
completion.

No local unit-test run: `nvidia-modelopt` isn't installed in my
checkout, so `tests/conftest.py` fails to import. Relying on CI.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — `config_overrides` defaults
to `None`; the `merge_lora` loader swap keeps the same class, dtype and
device_map for plain LLMs.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A — no new
dependency; `transformers_cosmos3` is an optional import guarded by
`contextlib.suppress`.
- Did you write any new necessary tests?: ❌ — exercising these paths
needs a checkpoint with a nested `text_config`, which the unit suite has
no fixture for. Happy to add one if a reviewer can point me at a small
suitable model.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
❌ — can add a *Speculative Decoding* entry for the `merge_lora` VLM fix
if you consider it changelog-worthy.
- Did you get Claude approval on this PR?: ❌ — not yet run.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added JSON-based model configuration overrides across speculative
decoding, training, validation, export, and LoRA workflows.
  - Overrides can update primary model and text configuration settings.
- Expanded support for vision-language models and Cosmos3 Omni
checkpoints.

- **Bug Fixes**
- Improved configuration handling for offline loading and
checkpoint-based initialization.
- Restored draft-model precision during checkpoint loading and model
conversion.
- Added validation for malformed, unsupported, and non-finite override
values.
- Standardized configuration override guidance across command-line
workflows.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix

Fix FP8 ONNX export for BF16 models during real-weight compression
without changing the public API or the `weights_dtype="fp32"` default.

The FP8 exporter preserves BF16 initializer bits when bridging
GraphSurgeon NumPy arrays to Torch, widens BF16 values exactly to FP32
for normalization, and leaves existing FP16/FP32 handling unchanged.
Conv scales and dequantized outputs retain the source dtype, and scales
round upward when needed so serialized values cannot cause FP8 overflow.

`weights_dtype="bf16"` is accepted as a no-op only for FP8-only models
whose floating parameters are all BF16. Registered buffers do not affect
this weight-focused decision and may preserve higher-precision regions
in the exported graph. Unsupported BF16 FP8-to-FP16 and FP32 or
mixed-parameter-to-BF16 conversions are rejected with `ValueError`
before temporary export paths are created. A narrow GraphSurgeon fix
preserves integer BF16 value-info dtypes.

### Usage

```python
onnx_bytes, metadata = get_onnx_bytes_and_metadata(
    quantized_fp8_model,
    (sample_input,),
    weights_dtype="bf16",
    onnx_opset=23,
)
```

### Testing

- Seven focused CPU regressions passed: BF16 QDQ compression and integer
dtype handling, BF16-to-BF16 and FP32-to-FP16 Conv/Linear export, and
four unsupported-conversion cases.
- QDQ utilities: 31 passed; pytest 2.25s, wall 29.88s.
- FP8 MHA exporter: 6 passed; pytest 2.05s, wall 32.71s.
- Torch deploy utilities: 51 passed; pytest 8.94s, wall 25.74s.
- Torch ONNX CPU export: 36 passed; pytest 4.85s, wall 32.38s.
- Changed-file pre-commit hooks: all passed; wall 8.07s.
- Exact-head FP8 BF16 GPU workflow at `f21d62a`: exit code 0; ONNX
checker passed; 6 FP8 initializers, 3 native `DequantizeLinear` nodes,
and 12 BF16 initializers.
- Refreshed GitHub CI at `f21d62a`: 50 passed and 1 skipped. Unit, GPU,
and regression required aggregates and Codecov passed. Two ONNX example
leaves failed because the runner could not load a cuDNN sublibrary;
their dependent example aggregate consequently failed.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅

### Additional Information

- TODO: Deliver authoritative `native`/FP32/FP16/BF16 ONNX export across
all quantized formats in follow-up pull requests.

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: export PTS/finetuned model to Hugging Face checkpoint,
then replace trtllm-build with trtllm-serve
Renamed export_trtllm_ckpt.py to export_hf_ckpt.py.
Replaced the legacy export_tensorrt_llm_checkpoint() flow with
export_hf_checkpoint().
Fix bug: 5823190

<!-- Details about the change. -->

```
python examples/llm_sparsity/weight_sparsity/hf_pts.py --model_name_or_path Llama-3.1-8B-Instruct --device cuda --model_max_length 1024 --dtype fp16 --sparsity_fmt sparsegpt  --calib_size 128  --output_dir Llama-3.1-8B-Instruct_pts

python examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py --model_name_or_path Llama-3.1-8B-Instruct --model_max_length 1024 --dtype fp16 --modelopt_restore_path Llama-3.1-8B-Instruct_pts/pts_modelopt_state.pth --output_dir Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts

trtllm-serve Llama-3.1-8B-Instruct_pts/trtllm/ckpt_pts \
    --tp_size 1 \
    --pp_size 1 \
    --host 0.0.0.0 \
    --port 8000

```

PTS and SAT tested

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

* **Documentation**
* Updated sparsity example instructions to export Hugging Face
checkpoints and serve models with `trtllm-serve`.
* Documented tensor and pipeline parallelism, host and port settings,
and the OpenAI-compatible chat completions endpoint.
  * Corrected the PTS model restoration path.

* **Bug Fixes**
  * Model export now saves the tokenizer alongside the checkpoint.
  * Model length configuration is interpreted as an integer.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change:  Bug fix:6701737

The ONNX deployment path assumed that ModelProto.ByteSize() would always
return a valid size. With newer protobuf versions, querying the size of
a model exceeding the protobuf serialization limit can itself raise
EncodeError: Failed to serialize proto.

Replaced both direct size checks with the existing
is_model_too_large_for_protobuf() helper. This helper handles size-query
failures conservatively and checks the protobuf size limit:

Shape inference now selects the external-data/file-based path when
ByteSize() fails or the model is too large.
Metadata creation uses the same safe check instead of raising another
serialization error.
The unused TWO_GB constant was also removed.

### Usage

```
python examples/diffusers/quantization/diffusion_trt.py --model flux-dev --benchmark --skip-image
```

### Testing
the above test case pass on B100

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?:  N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved ONNX model size detection during shape inference and export
processing.
* Ensured large models consistently use the appropriate external-data
handling path.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Noey Yang <174223378+noeyy-mino@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: documentation

Expose `mse_calibrate` through `model_calib.__all__` so Sphinx
autosummary includes the existing MSE calibration API on the generated
`model_calib` reference page.

The documentation configuration honors each module's curated `__all__`
surface. Although `mse_calibrate` was implemented and used by the
calibration dispatcher, it was missing from that surface and was
therefore filtered out during API generation.

### Usage

```python
from modelopt.torch.quantization.model_calib import mse_calibrate
```

### Testing

- `pre-commit run --files modelopt/torch/quantization/model_calib.py`
- `git diff --check -- modelopt/torch/quantization/model_calib.py`
- Generated the recursive autosummary API tree using the repository
Sphinx configuration and module template.
- Verified the generated RST contains `mse_calibrate`.
- Rendered the focused module page and verified its function-table link,
anchor, signature, and docstring.

The canonical full build was unavailable in the active environment
because the configured `shibuya` theme is not installed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — the existing
documentation build directly exercises this declarative autosummary
contract.
- Did you update Changelog?: N/A — this is a documentation-visibility
repair for an existing API.
- Did you get Claude approval on this PR?: N/A

### Additional Information

No source implementation or calibration behavior changed.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Made the MSE calibration capability publicly available for
quantization workflows.

- **Chores**
- Increased documentation build time limits to improve reliability for
longer-running builds.
- Increased multi-version test time limits to better accommodate
extended test runs.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: realAsma <akuriparambi@nvidia.com>
Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix for
#2209

Fix TEGroupedMLP per-expert weight quantizer checkpoint resharding.

`TEGroupedMLP` now saves its per-expert quantizer state as singleton
local shards, allowing the distributed checkpoint format to retain each
expert's global identity. Restore also initializes scalar `_amax`
placeholders after ModelOpt extra-state restoration so distributed
checkpoint loading can populate quantizer state for experts that move
between ranks.

This fixes restoring quantized TEGroupedMLP checkpoints across
expert-parallel and tensor-parallel topology changes.

Previously there was a bug that had two parts
1. TEGroupedMLP did not mark its per-expert quantizer state as
singleton_local_shards. That meant the scalar
weight_quantizer.<expert>._amax state was not saved with the same
globally unique expert identity as the grouped-expert weights, so DCP
could not reliably redistribute it across EP layouts.

2. During restore, ModelOpt’s extra-state restoration can leave _amax
absent for experts that were not local on the checkpoint’s saving rank.
The subsequent distributed checkpoint load then had no destination
tensor to populate.

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing

- `ruff format`, `ruff check`, `mypy`, `bandit`, and repository
pre-commit hooks
- Focused GPU regression:

  ```bash
python3 -m pytest
tests/gpu_megatron/torch/quantization/plugins/test_megatron.py \
    -k te_grouped_sharded_state_dict_reshard -v

Replaced the prior metadata-only TEGroupedMLP sharded-state test with an
end-to-end distributed-checkpoint save/restore regression test.

The new test:
- Quantizes a TEGroupedMLP with per-expert NVFP4 weight quantizers.
- Assigns each local expert a distinct, deterministic `_amax` based on
its global expert index.
- Saves both the model distributed checkpoint and sharded ModelOpt
state.
- Rebuilds the model under a different TP/EP topology.
- Restores ModelOpt state, loads the distributed checkpoint, and
verifies each target-local expert received the expected global-expert
`_amax`.

The parameterized test covers:
- EP=2 -> EP=1
- EP=1 -> EP=2
- TP=1 -> TP=2
- TP=2 -> TP=1
### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Improved checkpoint restoration for grouped quantizers by initializing
missing quantization statistics with compatible shapes.
- Improved restoration across supported grouped quantizer
configurations, including sequential groups and parallel checkpoint
layouts.
- Extra module state is now finalized through supported post-load
callbacks when available.
- Preserved populated quantized output-layer state during checkpoint
operations while removing empty placeholders.

- **Tests**
- Expanded checkpoint resharding coverage across tensor- and
expert-parallel configurations.
- Added coverage for disabled, dynamic, and other grouped quantizer
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Jenny Chen <jennifchen@nvidia.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…uantizer (#2203)

### What does this PR do?

Type of change: New feature (fail-fast guard; behavior change on a
previously silent path)

A `quant_cfg` whose module patterns don't match the model is not an
error to `set_quantizer_by_cfg` — every pattern simply matches nothing.
The run then calibrates, exports, and hands back a checkpoint that is
silently unquantized:

```json
{"quantization": {"quant_algo": null, "kv_cache_quant_algo": "FP8", "quantized_layers": {}}}
```

Nothing in the run says so. It has only ever been caught by someone
reading the exported `hf_quant_config.json` afterwards — most recently
on Step-3.7 ([NVBug 6518665](https://nvbugspro.nvidia.com/bug/6518665),
after a full 8×B200 calibration), and before that on MiniMax-M3, where
fused-expert detection skipped the experts and an experts-only recipe
matched nothing.

`mtq.quantize` now compares the config's intent against the outcome and
raises **before calibration**:

```
RuntimeError: The quantization config asks for weight quantization but no weight quantizer was
enabled, so nothing would be quantized (3 quantizer(s) inserted). These patterns matched no
weight quantizer:
  *.experts.*weight_quantizer
Either the patterns do not match this architecture's module names (check the model-specific
recipes under modelopt_recipes/huggingface/<model_type>/), or the modules holding the weights
were never converted to quantized modules (an unsupported custom module, e.g. a
trust_remote_code MoE layout).
```

Scoped to avoid false positives:

- **Only configs that ask for weight quantization** are checked (an
entry with `enable` and `weight_quantizer` in its pattern), so
activation-only and KV-cache-only configs are unaffected.
- **Intent is read from each pattern's final entry**, since `quant_cfg`
entries apply in order: a pattern that is enabled and then disabled
later asks for nothing by the end.
- **Configs refining an already-quantized model** (weight quantizers
enabled by an earlier `mtq.quantize`) are left alone.

Matching goes through `conversion._match_quantizer` — the same matcher
`set_quantizer_by_cfg` used to apply the config — so "did this pattern
match anything?" is answered exactly as the applying code would. A local
`fnmatch` diverges on the two cases that matcher handles:
`SequentialQuantizer` modules (W4A8-style list-valued `cfg`) and
fused-experts names (`..._weight_quantizers.0` normalizing to
`..._weight_quantizer`).

### Usage

No API change. A config that would previously have produced an
unquantized checkpoint now raises:

```python
mtq.quantize(model, {"quant_cfg": [
    {"quantizer_name": "*", "enable": False},
    {"quantizer_name": "*.experts.*weight_quantizer", "cfg": {"num_bits": 8, "axis": 0}},
]}, forward_loop)   # RuntimeError if the model has no `experts` modules
```

### Testing

Seven tests in `tests/unit/torch/quantization/test_quantize_cpu.py`, one
per branch of the guard: patterns matching nothing raise; an
activation-only config still runs; weight patterns disabled by a later
entry still run; enabled-then-retracted patterns still run;
`SequentialQuantizer` (list-valued `cfg`) and fused-experts quantizer
names count as matched; and the already-quantized refinement path is
exercised. Each was checked to be non-vacuous by removing the
corresponding branch and confirming exactly that test fails.

**One existing test changed.**
`tests/gpu/torch/export/test_fsdp2_export.py` parametrized over
`NVFP4_MLP_ONLY_CFG`, but its `SmallQKVModel` has no MLP — so that case
ran the FSDP2 paths against an *unquantized* model, and the new guard
reported it (4 GPU failures on the first CI run, all `quant_config6`;
`NVFP4_OMLP_ONLY_CFG` passed because that model does have `o_proj`). The
parametrization is dropped with a comment; `NVFP4_OMLP_ONLY_CFG` keeps
the scoped-recipe coverage. **If reviewers would rather not change that
test's meaning, the alternative is to downgrade the guard to a warning —
flagging it explicitly as a decision.**

I also swept every shipped `mtq.*_CFG` against `SmallQKVModel`: only the
four MLP/experts-scoped configs raise, and the other three
(`NVFP4_EXPERTS_ONLY_CFG`, `MXFP4_MLP_WEIGHT_ONLY_CFG`,
`NVFP4_MLP_WEIGHT_ONLY_CFG`) are used elsewhere only against real MoE
models (Qwen3-MoE, gpt-oss), so no other test is affected.

Ran locally after rebasing onto current `main` (torch 2.11, transformers
5.5.4): `tests/unit/torch/quantization` + `tests/unit/recipe` — 1271
passed, 7 skipped. Full `tests/unit` (minus onnx, and puzzletron which
needs `hydra`): 2676 passed, with 4 pre-existing
`test_quant_aware_conversion.py` failures that reproduce unchanged on
clean `main`. GPU tests were not run locally (no suitable GPU); the
FSDP2 change above is reasoned from the CI failure, not re-run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ❌ — deliberately. A config that
previously produced a `quant_algo: null` checkpoint now raises. Any such
run was already not doing what it claimed; the three scoping rules above
keep intentional non-weight quantization working.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ (Backward Breaking Changes)
- Did you get Claude approval on this PR?: ❌

### Additional Information

Pairs with #2202 (PTQ support for Step-3.7 MoE checkpoints), which fixes
the specific model that motivated this. Independent branches; either can
merge first.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Quantization now detects enabled weight-quantization patterns that do
not apply to any model weights and reports a clear validation error
before calibration.
- Broad wildcard patterns and nested quantizers are now handled
correctly.
- Overlapping patterns respect the final matching setting, including
later disabling rules.
- Existing quantized models can be refined using the parsed
configuration.
- Activation-only and explicitly disabled weight-quantization
configurations remain supported.
- Pipeline-parallel stages without targeted weights can bypass this
validation when configured to do so.

- **Documentation**
- Documented the process-wide override for bypassing unmatched
weight-quantizer validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: documentation

Align the ONNX PTQ README, guide, and executable example with the
implemented contracts:

- use the canonical `--calibration_data_path` CLI option;
- load `.npy` calibration data before passing it to the Python API;
- document the supported Autotune modes and calibration methods;
- correct the minimum opsets to INT8 19, FP8 19, and INT4 21; and
- describe the no-data fallback as random calibration inputs.

This also removes an inaccurate source comment without changing runtime
behavior.

### Usage

```bash
python -m modelopt.onnx.quantization \
    --onnx_path=model.onnx \
    --quantize_mode=int8 \
    --calibration_data_path=calib.npy \
    --output_path=model.quant.onnx
```

### Testing

- `pre-commit run --files docs/source/guides/_onnx_quantization.rst
examples/onnx_ptq/README.md modelopt/onnx/quantization/quantize.py
tests/examples/test_onnx_ptq.sh`
- `bash -n tests/examples/test_onnx_ptq.sh`
- `CUDA_VISIBLE_DEVICES="" python -m pytest -o addopts="" -p
no:cacheprovider --confcutdir=tests/unit/onnx/quantization -q
tests/unit/onnx/quantization/test_autotune_quantization_integration.py`
(4 passed)
- `nox -s docs` (passed; Sphinx built 881 HTML files)
- Focused before/after contract probe covering the documented CLI
option, API data type, Autotune modes and methods, opset minimums, and
random-input wording

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

Tracking: [6701308]

> 🤖 _Generated by Codex (AI agent)._

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Documentation**
- Clarified that random calibration inputs are used when no calibration
dataset is provided.
- Updated ONNX post-training quantization examples with minimum opset
requirements and the `calibration_data_path` argument.
- Clarified Autotune support for FP8 and INT8 calibration methods using
`max` or `entropy`.

- **Tests**
- Updated quantization command examples to use the current calibration
data path option.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

FP16 entropy calibration can fail for sufficiently narrow activation
ranges because NumPy may construct the histogram bin edges at FP16
precision. NumPy 2.2 and later reject the resulting collapsed bin
spacing, while earlier versions can silently return invalid,
non-monotonic edges. The same precision issue can recur when ONNX
Runtime merges later calibration batches.

Losslessly widen FP16 activation values to FP32 while calculating and
merging histograms in both ONNX entropy calibration paths, then restore
the source dtype at the calibration-to-quantization boundary. This keeps
the histogram bins stable without changing FP16 Q/DQ scale or graph
dtype semantics. FP32 inputs and public APIs are unchanged.

For full-range FP16 activations, ONNX Runtime can overflow while
subtracting FP16 calibration endpoints before it widens the result.
Retry only a non-finite FP16 scale calculation with FP32 endpoints, then
cast the finite scale back to FP16. Existing finite FP16 calculations
and all non-FP16 calculations continue to use ONNX Runtime's original
result.

The fallback intentionally patches only the `qdq_quantizer` binding used
for calibrated activation ranges. Initializer and weight quantization
continue to use ONNX Runtime's existing `quant_utils` path unchanged;
full-range FP16 weight scaling is outside this calibration fix.

The regression tests exercise both collectors across initial collection,
an equal-range merge, and an expanding-range merge. They also verify the
internal FP32 histogram and external FP16 calibration-range contract,
including finite saturation when restoring sanitized values. A real
entropy calibration test covers full-range FP16 values and verifies
finite FP16 Q/DQ scales and a loadable ONNX Runtime graph. The AutoCast
integration verifies the same FP16 scale-type contract through the
public quantization path.

N/A — no API or usage change.

All tests ran with CUDA hidden.

- NumPy 1.26.4: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.2.3: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- NumPy 2.3.5: affected histogram, AutoCast, and ORT-patching modules:
`39 passed`.
- Public INT8 entropy quantization with full-range FP16 calibration
data: finite FP16 Q/DQ scales, full ONNX check passed, and the CPU ONNX
Runtime session loaded.
- Changed-file pre-commit hooks: passed.

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌

Follow-up to #1558.

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix

AutoQuantize can measure a group of quantized expert layers at their
enclosing MLP output. That enclosing module is often a plain PyTorch
container and does not carry distributed-group information, so its
sensitivity score was not combined across data- or expert-parallel
workers.

This PR obtains the distributed groups from the quantized layers when
the scoring module does not provide them. It also preserves construction
order for quantized modules, scoring modules, and their registered
hyperparameters so every worker accumulates scores in the same order.

The temporary state needed by backward-based scoring is now managed by
one shared session. The session installs and removes forward patches and
invocation-specific output-gradient hooks, controls parameter gradients,
and restores the active quantization recipes even when scoring raises an
exception. Scoring methods remain responsible for their own score
calculation.

### Usage

N/A — this fixes existing AutoQuantize behavior and does not add an API
or flag.

### Testing

- `pre-commit run --files modelopt/torch/quantization/algorithms.py
tests/unit/torch/quantization/test_autoquant.py`
- `pytest -q tests/unit/torch/quantization/test_autoquant.py` — 102
passed
- Added a real two-rank gradient AutoQuantize test covering MoE experts
scored at an enclosing MLP.
- Added regressions for deterministic hyperparameter registration,
per-invocation replay for reused score modules, exact
`forward`-attribute restoration, partial setup rollback, and cleanup
after a scoring failure.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors).

- Is this change backward compatible?: ✅ — no API or checkpoint format
changes; distributed sensitivity values now include the missing
reduction.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A — no new feature, deprecation, breaking change, or critical
release-note item.
- Did you get Claude approval on this PR?: ❌ — pending review.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
* Improved quantization scoring consistency through deterministic
ordering and invocation handling.
* Added more reliable distributed score aggregation, including support
for mixture-of-experts models.
* Improved gradient-based scoring for repeated evaluations, tuple
outputs, and checkpoint-compatible workflows.
* Ensured model behavior and scoring state are restored after successful
or failed evaluations.
  * Avoided unnecessary output replay when gradients are not required.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Joshua Hill <joshua.hill@baseten.co>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: documentation

`general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml` appears in the
shipped-recipes
table in `modelopt_recipes/ptq.md`, but the **Calibration variants**
section —
which documents `max`, `mse`, `input_scale1`, `gptq`, and the
`layerwise`
variants — had no entry for it. Someone scanning that section for "which
calibration do I pick when NVFP4 W4A4 regresses?" only found `mse`,
which
searches **weight** scales and so cannot help when the loss comes from
activation clipping.

This adds the missing entry: the scale formula
(`amax = max(rho * anchor, upper)`) and its defaults, the fact that it
costs one
calibration pass and exports a standard NVFP4 checkpoint with coverage
identical
to `nvfp4_default-kv_fp8_cast`, and the symptoms that should route you
here
rather than to a weight-side calibration — an A16 ablation clears the
regression
while `mse` does not, the symptom is behavioral (verbose or runaway
generations,
hitting the generation cap) rather than a flat score drop, inference
contexts run
longer than the calibration set, or a few rare blocks dominate the
activation
error. MoE experts-only scopes are called out as the common case.

It also extends step 3 of **Choosing a general recipe** so the
escalation path
reads `mse` first, then `nvfp4_act_headroom` when the evidence points at
activations rather than weights.

**Evidence.** The guidance comes from a GLM-5.3-Flash NVFP4 experts-only
W4A4
root-cause study on SciCode (temperature 1.0), which established
causally that
activation quantization at the routed-expert `down_proj` input drove a
large
generation-length blow-up. Swapping `max` for `nvfp4_act_headroom` cut
the median
generation-length regression versus source from +38% to +19% and the
mean from
+19% to +4%, with no capped generations. The entry states plainly that
this was
the best strict-W4A4 result in that study but still missed the p50/p75
near-lossless gate, so headroom is presented as a strong first lever for
activation-driven regressions rather than a guaranteed fix, with a note
that
`rho` should be swept.

### Usage

No API or recipe change; the recipe already ships. This PR only
documents when
to select it over plain `max`:

```python
from modelopt.recipe import load_recipe

cfg = load_recipe("general/ptq/nvfp4_act_headroom-kv_fp8_cast")
```

### Testing

Docs-only change; no code paths touched, so no new or updated tests.

- `pre-commit run --files modelopt_recipes/ptq.md` — all applicable
hooks pass,
  including `markdownlint-cli2` and `check-modelopt-recipes`.
- Re-read the rendered section to confirm the new bullet nests correctly
in the
existing `Calibration variants` list and that surrounding entries are
unchanged.
- Cross-checked every claim against the implementation
(`modelopt/torch/quantization/calib/nvfp4_act_headroom.py`), the recipe
YAML,
  and the existing `CHANGELOG.rst` entry, so the documented defaults
  (`anchor_percentile=1`, `upper_percentile=99.99`, `rho=16384`) and the
  NVFP4-input-quantizer-only scope match the code.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A <!-- docs-only; the
algorithm's tests already live in
tests/unit/torch/quantization/test_nvfp4_act_headroom.py -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A <!-- nvfp4_act_headroom already has a CHANGELOG entry from the PR
that added it; a docs-only follow-up is not changelog-worthy. -->
- Did you get Claude approval on this PR?: ❌ <!-- not run; docs-only
change -->

### Additional Information

The entry deliberately does not sell this on accuracy: in that study the
quantized subtask accuracy (56.80%) was *above* the source checkpoint
(51.18%),
so what headroom recovered was generation-length behavior, not accuracy.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added an NVFP4 activation headroom calibration option with
configurable percentile-based scaling.
* Supports standard NVFP4 checkpoint export with a single calibration
pass.
* Applies to dynamic-block NVFP4 activation quantizers while keeping
weight-scale configuration independent.
* Added guidance for addressing activation-related W4A4 accuracy
regressions, including calibration coverage and recipe selection.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chenjie Luo <chenjiel@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Chad's Agent

Type of change: documentation.

Replace legacy TensorRT checkpoint export, support matrix, and
engine-build instructions with `export_hf_checkpoint` and TensorRT-LLM's
PyTorch backend. Preserve the existing 0.48.0 deprecation / 0.49.0
removal notice and page URL. Update the customized-model guide and
deployment skill to match.

Follow the linked unified HF export guide. No API changes.

- `git diff --check` passed.
- `uvx pre-commit run --files docs/source/deployment/1_tensorrt_llm.rst
docs/source/guides/_customized_model_quantization.rst
plugins/modelopt/skills/deployment/references/trtllm.md` passed.
- Full Sphinx build delegated to the Docs workflow; preview expected
after deployment.

- Is this change backward compatible?: ✅ Documentation only; page URL
retained.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A — documentation only.
- Did you update Changelog?: N/A — guidance correction; no new API
deprecation.
- Did you get Claude approval on this PR?: ❌ Not requested yet.

Removes instructions for the TensorRT backend that current TensorRT-LLM
releases no longer support.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **Documentation**
- Updated TensorRT-LLM deployment guidance to use `export_hf_checkpoint`
with the PyTorch backend.
- Clarified that this workflow does not require TensorRT engine
construction.
- Updated DBRX customization instructions for exporting and deploying
quantized models.
- Added TensorRT-LLM version requirements and links to unified Hugging
Face deployment guidance.
- Removed guidance for the legacy TensorRT-LLM checkpoint exporter and
outdated troubleshooting steps.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

Fail fast when AutoQuantize receives non-finite output gradients, before
accumulating sensitivity scores. The error names the affected module and
suggests checking the model, data, and loss; it also identifies cuDNN
SDPA backward on fully masked rows as one possible cause and gives an
explicit retry workaround.

Unlike the earlier revision, this does not disable cuDNN or change any
attention backend settings. Invalid gradients are not zeroed or ignored,
and there is no automatic retry.

No API or recipe changes. For the reproduced cuDNN failure, the caller
can explicitly set `torch.backends.cuda.enable_cudnn_sdp(False)` before
a fresh AutoQuantize run.

- AutoQuantize unit suite: **110 passed**. Coverage includes NaN and
positive/negative infinity, module diagnostics, preventing invalid score
accumulation, model-state cleanup, and unchanged SDPA backend settings.
- Real-model E2E on **four GB300 GPUs**, Qwen/Qwen3.6-35B-A3B, main
`8025a3dc5481129aa21fef99cb13a879e1b5847e` plus this patch, batch size
8, 512 calibration samples, and
`w4a16_nvfp4_fp8_at_6p0bits-active_moe.yaml`:
- Default backend: the new diagnostic fired at
`model.language_model.layers.39.self_attn.q_proj`; cuDNN remained
enabled and no quantized model was exported. Expected-error check
passed.
- Explicit cuDNN-disabled fresh run: both 64-batch calibration passes,
all 16 scoring batches, optimization at **5.99 effective bits**, and
checkpoint export completed with exit code 0. Verified all three indexed
safetensors shards and quantization configuration; the index contains
93,563 tensors.
- All applicable pre-commit checks and `git diff --check` passed.

Contributor guidelines and security guidance followed; commits are
signed and signed off.

- Is this change backward compatible?: Yes; finite-gradient behavior and
backend settings are unchanged.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A; neither
added.
- Did you write any new necessary tests?: Yes.
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
Yes, 0.48.0 bug fixes.
- Did you get Claude approval on this PR?: No; awaiting review.

This improves error reporting rather than fixing the upstream cuDNN
kernel. Blackwell-specificity is not established. Checkpoint
deployment/reload was not tested.

A calibration-only checkpoint-resume attempt completed scoring but
encountered a separate `candidate_stats` KeyError. That issue is outside
this patch; the successful export validation above used a fresh run
without search-state resume.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

* **Bug Fixes**
* AutoQuantize now fails fast when output gradients contain non-finite
values, with an actionable error identifying the affected module.
* Attention backend settings are preserved and restored after successful
runs and failures.
* Model state is restored when sensitivity scoring encounters an error
during setup or execution.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: weimingc <17592131+meenchen@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…tate dump (#2410)

### What does this PR do?

Type of change: Bug fix

Fixes `nvbugs/6753684`, filed against #2080 by the ModelOpt QA Sentinel.

The vLLM offline hidden-state dump rejected `--aux-layers eagle` — **the
flag's own default** — so the documented invocation aborted before
writing any state:

```
File "collect_hidden_states/compute_hidden_states_vllm.py", line 76, in _resolve_aux_layers_standalone
    ids = sorted({int(t) for t in aux_layers.split(',') if t.strip()})
ValueError: invalid literal for int() with base 10: 'eagle'
```

**Root cause.** `compute_hidden_states_vllm.py` runs in a stock vLLM
container, where importing `modelopt.torch` fails (the full init chain
pulls in omegaconf and friends). It therefore carries
`_resolve_aux_layers_standalone`, a local copy of the preset logic in
`common.resolve_aux_layers`. That copy implemented the `dflash` preset
and explicit id lists, but never `eagle` — while `add_aux_layers_args`
defaults to `eagle`. The HF and TRT-LLM dumps call the shared helper and
were unaffected; only the vLLM path forked, and nothing compared the
fork against its source.

This PR resolves `eagle` inline, mirroring
`hf_eagle.default_eagle_aux_layer_ids`.

It also fixes a second defect the bug exposes: the function already had
a message naming the accepted values, but it was unreachable, because
`int()` raised first. An unrecognised preset now reports what it accepts
instead of surfacing the raw `int()` error — which is what made the
original failure opaque.

### Usage

The previously-broken documented invocation now works:

```bash
cd examples/speculative_decoding
python collect_hidden_states/compute_hidden_states_vllm.py \
    --model Qwen/Qwen2.5-0.5B-Instruct \
    --input-data ../dataset/synthetic_conversations_1k.jsonl \
    --output-dir /tmp/hs_vllm \
    --max-seq-len 512 --tp 1
```

`--aux-layers dflash` and explicit lists such as `--aux-layers 2,5,8`
are unchanged.

### Testing

Added `tests/unit/examples/test_vllm_hidden_states_aux_layers.py`, which
pins the standalone copy to the shared implementation it mirrors:

- `eagle` matches `hf_eagle.default_eagle_aux_layer_ids` across layer
counts 4, 6, 8, 12, 24, 28, 32, 36, 48, 52, 61, 80 — deliberately
including counts small enough that the `max(0, ...)` clamps collapse ids
together.
- A named regression case for `nvbugs/6753684`.
- `dflash` and explicit-list behaviour unchanged.
- Unknown specs (`bogus`, `EAGLE3`, `eagle3`, empty) raise the
actionable message.
- Out-of-range ids still rejected.

Divergence here is silent — the dump would write plausible-looking
hidden states from the *wrong* layers, surfacing much later as a poor
acceptance rate. Hence pinning to the reference rather than asserting
hardcoded lists alone.

All 20 assertions verified and every pre-commit hook passes (`ruff`,
`mypy`, `bandit`, RST lint, license headers).

One caveat worth stating plainly: **pytest could not be run locally.**
`tests/unit/conftest.py` imports `modelopt.torch.utils.distributed`,
which needs `CPUOffloadPolicy` from `torch.distributed.fsdp` — absent in
this machine's torch. Each assertion was executed directly against the
real module instead, but CI is the first genuine pytest run.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅ — strictly widens accepted
input; `dflash` and explicit lists behave identically.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — bug fix for a defect present in a previous release.
- Did you get Claude approval on this PR?: ❌ — not yet run.

### Additional Information

The underlying fragility is the duplicated implementation, not this one
missing branch. The function's own `TODO: drop this once
common.resolve_aux_layers is decoupled from the heavy modelopt.torch
import chain` is the real fix; the new test narrows the gap but does not
close it. Worth tracking separately if the vLLM dump is expected to keep
pace with new presets.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Bug Fixes**
  * Fixed `--aux-layers eagle` for vLLM offline hidden-state collection.
* Added support for the documented `eagle` preset alongside `dflash` and
explicit layer IDs.
  * Improved invalid-option errors to clearly list accepted formats.
* Rejects `dflash` configurations when the target model has too few
layers.
  * Continues rejecting layer IDs outside the model’s available range.
* **Documentation**
  * Added a v0.48.0 changelog entry for the fix.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix

Fix fakequant calibration for hybrid attention/Mamba models, including
NVIDIA Nemotron-3-Nano, on vLLM 0.26 and 0.28.

The manual calibration scheduler path previously submitted requests with
empty KV-cache block tables. Hybrid models require scheduler-compatible
cache state during prefill; on current vLLM releases the empty tables
caused the Mamba state to use the reserved null block and calibration
activations became NaN. Request cleanup also no longer matched the vLLM
0.28 execution lifecycle, which could leave request-scoped state in the
persistent batch.

This PR:

- Allocates non-null scratch blocks for every KV-cache group using the
vLLM warmup reservation policy.
- Supports both the vLLM 0.28 reservation helper and the equivalent vLLM
0.26 calculation.
- Passes newly allocated blocks through `new_block_ids_to_zero` when
that scheduler field is available.
- Validates that the calibration batch fits in the configured cache and
reports how to reduce calibration demand if it does not.
- Cleans up calibration requests through a zero-token scheduler step on
current vLLM, with a direct cleanup fallback for older runners.
- Updates the example Dockerfile to default to vLLM 0.28.0 while
retaining vLLM 0.26.0 through `VLLM_VERSION`.
- Documents the validated Nemotron-3-Nano NVFP4 KV-cache workflow and
clarifies that reducing `--max-num-batched-tokens` is not required.

### Usage

Build the default vLLM 0.28.0 image:

```bash
docker build -f examples/vllm_serve/Dockerfile \
  -t vllm-modelopt:v0.28.0 .
```

Build with vLLM 0.26.0:

```bash
docker build --build-arg VLLM_VERSION=0.26.0 \
  -f examples/vllm_serve/Dockerfile \
  -t vllm-modelopt:v0.26.0 .
```

Calibrate and serve Nemotron-3-Nano with NVFP4 KV-cache fakequant:

```bash
KV_QUANT_CFG=NVFP4_KV_CFG QUANT_CALIB_SIZE=512 \
  python examples/vllm_serve/vllm_serve_fakequant.py \
  <nemotron3_nano_model_path> \
  --trust-remote-code --enforce-eager -tp 8 \
  --max-model-len 8192 --host 0.0.0.0 --port 8000
```

### Testing

Validated on omniml-a0 with `NVIDIA-Nemotron-3-Nano-30B-A3B-BF16`,
tensor parallel size 8, `NVFP4_KV_CFG`, `QUANT_CALIB_SIZE=512`, and
`--max-model-len 8192`. No `--max-num-batched-tokens` override was used.

- vLLM 0.28.0:
  - All 512 calibration samples completed.
  - No NaNs or cache-cleanup warnings were observed.
  - The server started and `/health` passed.
- An OpenAI-compatible completion request returned coherent generated
text.
- vLLM 0.26.0:
- Repeated the same 512-sample TP8 calibration with the official
`vllm/vllm-openai:v0.26.0` image.
  - No NaNs were observed.
- The server started, passed `/health`, and returned coherent generated
text.
- Docker:
  - Built and verified the updated vLLM 0.28.0 image.
- Focused tests:
  - `tests/examples/vllm_serve/test_vllm_mlflow_utils.py`: 32 passed.
- Cleanup failure, missing legacy API, and legacy fallback tests: 5
passed on both vLLM 0.26.0 and 0.28.0.
- Repository hooks:
- Targeted pre-commit hooks for every changed Python, Markdown, and
Docker file: passed.
  - `git diff --check`: passed.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ — added focused coverage for
fail-closed cleanup, exception chaining, and the legacy cleanup
fallback; the full regression was also validated end to end.
- Did you update Changelog?: N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

The change is quantization-format agnostic. It corrects the calibration
scheduler and cache lifecycle rather than special-casing `NVFP4_KV_CFG`
or using an NVFP4 cast path.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added support for configuring the vLLM version through `VLLM_VERSION`,
with vLLM 0.28.0 as the default.
- Added calibration and serving guidance for hybrid attention/Mamba
models, including Nemotron 3 Nano with NVFP4 KV-cache fake quantization.

- **Bug Fixes**
  - Improved calibration block handling across supported vLLM versions.
- Improved calibration cleanup to preserve original errors and provide
reliable fallback behavior when standard cleanup is unavailable.

- **Documentation**
- Documented tested versions, direct installation commands, ModelOpt
setup, serving options, and guidance to avoid NaNs during batched
serving.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: Kinjal Patel <kinjalpravin@nvidia.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…ve (#2460)

### What does this PR do?

Type of change: Documentation

`dflash.yaml` tells the reader that chat templates live in
`modelopt_recipes/general/speculative_decoding/chat_templates/`. That
directory does not exist, and this comment is its only mention anywhere
in the tree:

```
$ git grep -n 'speculative_decoding/chat_templates'
modelopt_recipes/general/speculative_decoding/dflash.yaml:19: # Templates are in modelopt_recipes/general/speculative_decoding/chat_templates/
```

Templates actually sit beside each launcher example —
`tools/launcher/examples/Qwen/Qwen3-8B/chat_template_train.jinja`,
`.../MiniMax/MiniMax-M2.7-DFlash/chat_template_train.jinja`, and so on.

Split out of #2201, where this two-line comment was the only reason
`modelopt-recipes-codeowners` was a required reviewer on a
skills-documentation PR.

### Usage

No behaviour change — comment only.

### Testing

None needed; the file's only change is a YAML comment. `pre-commit`
passes.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **Documentation**
- Updated speculative decoding guidance to clarify where each model’s
chat template is maintained alongside its launcher example.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ye Yu <yeyu@nvidia.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

Preserves the public ONNX graph I/O types captured at the API boundary
when `PrecisionConverter` wires output casts. Type inference can change
the working graph's output declaration before conversion; consulting
that mutated declaration caused the required cast back to the original
public type to be discarded and metadata restoration to fail.

The converter now derives its I/O type map from the preserved boundary
metadata and uses that map when deciding whether a cast should become a
public graph output. A regression test covers an FP32 output whose
working declaration is inferred as FP16, and the changelog records the
corrected behavior.

```python
converted = convert_to_f16(model, keep_io_types=True)
```

- Ran `pytest tests/unit/onnx/autocast/test_precisionconverter.py` (186
passed).
- Ran `pytest tests/unit/onnx/autocast` (249 passed).
- Ran Ruff check and format validation on the changed Python files.
- Verified the original minimal end-to-end reproduction with Python 3.12
and TensorRT 10.16.1.11; conversion now completes without the
output-metadata restoration error.
- Verified the full CLI quantization path proceeds through the formerly
failing one-Q/DQ scheme and successfully benchmarks the generated
TensorRT engine.

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: N/A

Tracking: [6771663]

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **Bug Fixes**
- Fixed ONNX FP16 conversion to preserve public graph output types when
type inference changes internal declarations.
- Ensured output casts are inserted correctly when preserving
input/output types is enabled.

- **Tests**
- Added regression coverage confirming preserved output types, correct
cast insertion, and valid ONNX model generation.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Ajinkya Rasane <ajinkyaashwin@gmail.com>
Co-authored-by: Ajinkya Rasane <ajinkyaashwin@gmail.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

Adds a compact SDXL and SDXL-Turbo mixed-precision FP4 recipe:

- block-16 NVFP4 for non-QKV Linear/GEMM layers;
- FP8 for Conv2d layers;
- high-precision Q/K/V projection Linears to preserve TensorRT
horizontal fusion;
- optional FP8 MHA quantization.

For SDXL FP4 export, Conv2d quantizers export directly through the
shared FP8 custom-op path. The previous `generate_fp8_scales` plus
`convert_zp_fp8` INT8 zero-point workaround is removed. The graph then
uses the existing FP8 Q/DQ normalization and `NVFP4QuantExporter`
lowering, with opset 23 for FLOAT4 support. Flux FP8 export also saves
the graph returned by its RoPE weight conversion.

This PR also changes shared exporter behavior:

- `_fp8_quantize` refreshes ONNX shape/type inference after applying the
custom FP8 operator's uint8 output metadata, affecting all FP8 ONNX
exports through this symbolic.
- `_quantized_sdpa` derives `disable_fp8_mha` from the live Q/K/V
quantizer state instead of a restored private module flag.

Other model recipe configurations remain unchanged.

```bash
python quantize.py \
    --model sdxl-1.0 \
    --model-dtype Half \
    --trt-high-precision-dtype Half \
    --format fp4 \
    --block-size 16 \
    --batch-size 2 \
    --calib-size 128 \
    --n-steps 20 \
    --quantized-torch-ckpt-save-path ./sdxl-fp4 \
    --onnx-dir ./onnx-sdxl-fp4
```

- CPU-only focused and generic NVFP4 exporter tests: 44 passed in 4.35
seconds.
- Focused Flux returned-graph save test: 1 passed.
- Required Linux unit CI at `034fe23ec` passed with the `all` dependency
set, including `tests/unit/examples/test_diffusers_fp4.py`.
- Latest changed-file pre-commit checks: all passed.
- TensorRT 10.14 on a B200 GPU:
  - 302 native block-scaled NVFP4 GEMM tactics;
  - 38 native FP8 Conv tactics;
  - no FP4 Q/K/V projections;
  - all 11 FP16 Q/K/V projection-fusion groups preserved;
- three alternating batch-2 profiles measured 18.614 ms FP4 versus
20.028 ms FP16 median UNet latency, a 7.06% reduction.
- FP8 SDXL/SD3 ONNX-to-TensorRT end-to-end runs were not executed
because they require explicit approval. The existing end-to-end test
matrix now includes SD3 FP8 alongside SDXL FP8.

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ — no public API or CLI flags
change; the shared changes preserve the intended FP8 export and
attention behavior.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ — the shared NVFP4 opset, FP8 shape-inference, and Diffusers
attention-policy changes are recorded under bug fixes.
- Did you get Claude approval on this PR?: N/A

Tracking: [5565357]

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **New Features**
- Added SDXL support for mixed NVFP4/FP8 quantization, including
convolution and softmax handling.
- Added an SDXL quantization preset for streamlined post-training
quantization workflows.
- Expanded FP4 ONNX export support to Flux and SDXL, with improved
FP4/FP8 graph processing and export reliability.
- Added automatic quantization policy and format restoration from
checkpoints.

- **Documentation**
- Documented SDXL layer behavior, optional FP8 attention quantization,
and Blackwell/TensorRT requirements for FP4 and FP8 deployment.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

> 🤖 _Generated by Codex (AI agent)._

---------

Signed-off-by: ajrasane <131806219+ajrasane@users.noreply.github.com>
Co-authored-by: Codex <codex@openai.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix

<!-- Details about the change. -->

### Usage

```python
# Add a code snippet demonstrating how to use this
```

### Testing
<!-- Mention how have you tested your change if applicable. -->

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: ✅ / ❌ / N/A <!--- If ❌, explain
why. -->
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: ✅ / ❌ / N/A
<!--- Mandatory -->
- Did you write any new necessary tests?: ✅ / ❌ / N/A <!--- Mandatory
for new features or examples. -->
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅ / ❌ / N/A <!--- Very short summary of changes only for new features,
backward breaking changes, deprecations, or fixes for critical bugs
present in previous releases. -->
- Did you get Claude approval on this PR?: ✅ / ❌ / N/A <!--- Run
`/claude review`. NVIDIA org members can self-trigger for complex
changes; orthogonal to CodeRabbit. -->

### Additional Information
<!-- E.g. related issue. -->

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Documentation**
* Updated the NVIDIA Nemotron offline usage example to use the BF16
source-directory pipeline path.
  * Pipeline configuration remains unchanged.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Jennifer Chen <jennifchen@nvidia.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Fixes #1871

The pruning documentation is split between
`docs/source/guides/3_pruning.rst`
and `examples/pruning/README.md`, causing confusion about which is
authoritative.

The examples/pruning/README.md is the comprehensive, up-to-date
reference
covering Minitron, Puzzletron, FastNAS, support matrix, guidelines, and
distillation hyperparameters.

This PR adds a note to the RST guide making clear:
- The README is canonical for Minitron and Puzzletron (LLM/VLM pruning)
- The guide covers FastNAS for Computer Vision models

Signed-off-by: Diya <diyaismahil7@gmail.com>

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Documentation**
- Updated the pruning guide’s introductory content for clearer
separation of general guidance and the related Minitron/Puzzletron note.
- Clarified that the guide focuses on FastNAS pruning for computer
vision models.
- Added references to the Pruning README for Minitron and Puzzletron API
examples, support information, guidelines, and distillation
hyperparameters.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: didi <diyaismahil7@gmail.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…re (#2480)

Type of change: Bug fix

`post_quantize()` in `examples/hf_ptq/hf_ptq.py` ran the optional
post-quantization sanity-check `full_model.generate()` unguarded,
directly before `export_quantized()`. Any exception raised there aborted
the whole run and discarded a completed calibration without exporting a
checkpoint.

Root cause (traced from [NVBug
6752977](https://nvbugspro.nvidia.com/bug/6752977), DGX Spark GB10 /
DeepSeek-R1-Distill-Llama-8B / NVFP4): `get_model()` loads with
`device_map="auto"`, relying on `accelerate`'s
`infer_auto_device_map`/`get_max_memory()` to decide GPU vs. CPU
placement. On DGX Spark's unified-memory single-GPU host, that memory
probe under-reports GPU capacity, so part of even an 8B model can land
on CPU — and the existing fallback shrinks the GPU budget further (`*
gpu_mem_percentage`), compounding it. Calibration survives this because
it never invokes the real fake-quant kernel, but the post-PTQ sanity
`generate()` does, and NVFP4's dynamic-block-quantize op
(`modelopt/torch/quantization/tensor_quant.py`) hard-asserts
`amax.is_cuda` with no CPU fallback, so any CPU-offloaded layer crashes
there — after ~5.8 hours of calibration, before export.

This PR does not attempt to fix the underlying
`device_map`/memory-probing behavior (unverified without the actual
hardware/logs, which weren't reachable from this environment). Instead
it makes the failure mode safe: a failure in the optional sanity check
now only skips that check and warns, and export always proceeds,
regardless of why `generate()` failed.

No new API. Behavior change only: `examples/hf_ptq/hf_ptq.py` now
completes export even if the post-quantization sanity `generate()` call
raises.

- Added
`tests/examples/hf_ptq/test_hf_ptq_args.py::test_post_quantize_export_survives_a_failed_sanity_generate`,
which drives `post_quantize()` with a `full_model.generate()` that
raises and asserts `export_quantized()` still runs.
- Ran `pytest tests/examples/hf_ptq/test_hf_ptq_args.py` (48 passed).
- Ran `pre-commit` on the changed files (`ruff-format` reformatted line
wrapping only).

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ <!-- pending: run `/claude
review` -->

Fixes NVBug 6752977. Linked JIRA: OMNIML-5932.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

- **Bug Fixes**
- Quantized checkpoint export now continues when the optional
post-quantization generation check fails.
  - A warning is shown when the generation check cannot complete.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Zhiyu Cheng <zhiyuc@nvidia.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
…2452)

Type of change: Bug fix

Hybrid (e.g. Nemotron-H) checkpoints saved by the
`examples/megatron_bridge` scripts could not be reloaded or exported to
HuggingFace:

TypeError: MLPSubmodules.__init__() missing 2 required positional
arguments: 'linear_fc1' and 'linear_fc2'

**Root cause.** Megatron-LM's YAML writer represents a
`functools.partial` via `_partial_representer`, which passes each
keyword value through `represent_data`. A dataclass instance has no
representer, so it falls through to `_safe_object_representer`, which
emits only `{_target_, _call_}` and drops every field. The default
hybrid stack spec builds its dense-MLP and MoE layers as exactly such
partials, and `set_moe_expert_layout()` stored the *built* `ModuleSpec`
on the provider — which is serialized into every checkpoint's
`run_config.yaml`. So `MLPSubmodules` / `MoESubmodules` were written
with no fields at all. Both export paths hit it: `convert.sh` via
`from_auto_config`, and `export_distilled_megatron_to_hf.py` via
`export_ckpt → load_megatron_model`. Every hybrid provider is affected,
dense or MoE.

**Fix.** `set_moe_expert_layout()` stores a named, zero-argument
*factory* instead. The provider already calls a callable spec at build
time (`_resolve_hybrid_stack_spec`), so model construction is unchanged
— only the serialized form differs:

```yaml
  hybrid_stack_spec:
    _call_: false
    _target_: megatron.bridge.models.hybrid.hybrid_provider.transformer_engine_hybrid_stack_spec
```

The grouped-GEMM factory is Megatron-Bridge's own
`transformer_engine_hybrid_stack_spec`, so stock tooling
(`scripts/conversion/convert.sh`) resolves it without importing
ModelOpt.

**Known limitation — the two layouts are not symmetric.** The
SequentialMLP layout has no bridge-side equivalent (the upstream TE
hybrid spec hardcodes `TEGroupedMLP`), so it serializes a ModelOpt
target, which `instantiate` only accepts in a process that has imported
`mbridge.py` and thereby run `register_allowed_target_prefix`. A
SequentialMLP hybrid checkpoint therefore converts through the ModelOpt
entrypoints but not through stock `convert.sh`, where it fails on the
disallowed prefix instead of on `MLPSubmodules` — no regression, but
that path stays broken for this one layout. The reach is narrow:
`use_moe_grouped_gemm()` returns True for any architecture with a
grouped-expert export rule, NemotronH included, so SequentialMLP
requires an explicit `--no_moe_grouped_gemm`. Closing it properly needs
an upstream `moe_grouped_gemm`-aware factory in Megatron-Bridge.

Both spec builders also move out of `nas/plugins/megatron.py`, which
never used them, into a new `utils/plugins/megatron_layer_specs.py`
beside the other Megatron-Core-only helpers. Not into `mbridge.py`: that
module needs `megatron.bridge`, while `get_te_hybrid_stack_spec` is
reached by 16 test files through
`tests/_test_utils/torch/megatron/models.py`, which is bridge-free.

The underlying defect is upstream in
`megatron/training/config/yaml_utils.py`; this only stops ModelOpt from
stepping on it, so it is worth a separate Megatron-LM issue.

No API change — hybrid checkpoints saved after this fix convert with the
existing commands:

```bash
torchrun --nproc_per_node 1 examples/megatron_bridge/export_distilled_megatron_to_hf.py \
    --student_hf_path <student_hf_model_or_path> \
    --megatron_path   <distill_out>/checkpoints \
    --hf_export_path  <hf_out> \
    --export_iterations all
```

Verified in `nemo:26.08` (megatron-core 0.19.0) against a 30B-A3B
Nemotron-3.5-Lightning pruned+distilled run:

- **Round trip, both MoE layouts.** Ran `set_moe_expert_layout` on a
real `HybridModelProvider`, dumped it through `dump_dataclass_to_yaml`
(the writer used for `run_config.yaml`), reloaded via `instantiate`,
resolved. `moe_grouped_gemm=True` →
`TELayerNormColumnParallelLinear`/`TERowParallelLinear` +
`TEGroupedMLP`; `False` → same MLP + `SequentialMLP`. The field stays
callable after `finalize()` and `_resolve_hybrid_stack_spec()`, so a
saved config cannot regress.
- Applying the equivalent `run_config.yaml` fix to 32 iteration
checkpoints: all 32 rebuild the provider (52 layers, hidden 2304, 104
experts) with populated `MLPSubmodules` / `MoESubmodules`.
- **End-to-end exports**, 6 iterations, all rc 0, each producing exactly
the source model's 5139 tensor keys (0 missing, 0 extra), 9 shards /
41.5 GiB, all weights finite, drift from the base rising monotonically
with iteration (lm_head 0.030 → 0.092). Covered `convert.sh` CPU,
`convert.sh` GPU (4×GB300, TP=4), and
`export_distilled_megatron_to_hf.py`. Same iteration and wrapper: CPU
123 s vs GPU 134 s — GPU is not faster, since with TP=4 each rank still
builds 20.9 B of 22.3 B params and the cost is I/O plus CPU-side
conversion.
- `ruff check` / `ruff format --check` passed on the source files before
the module move.

**Not yet run:**
`tests/gpu_megatron/torch/utils/plugins/test_mbridge.py` (added here) —
the GPU allocation expired. It asserts the round-trip property verified
manually above, but its `HybridModelProvider(num_layers=2,
hidden_size=64, num_attention_heads=4)` construction is unverified. The
module move is verified only by reference grep and syntax check, so
please also run one test that uses
`tests/_test_utils/torch/megatron/models.py`. `pre-commit` was not run
either (unavailable in the environment used).

- Is this change backward compatible?: ✅ behavior; note
`get_te_hybrid_stack_spec` moved module
(`modelopt.torch.nas.plugins.megatron` →
`modelopt.torch.utils.plugins.megatron_layer_specs`), and a checkpoint
from 0.46.1/0.47.0 needs the `run_config.yaml` edit described in the
changelog.
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅ (added, not yet executed —
see Testing)
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ❌ — will run `/claude review`
before marking ready.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

* **New Features**
* Hybrid checkpoints now record complete layer specifications in
`run_config.yaml`.
  * Recorded specifications support conversion to Hugging Face format.
* Hybrid MoE configurations support grouped-GEMM and sequential-MLP
modes.
* Configuration-based reconstruction preserves the selected MoE layout.

* **Compatibility**
* Checkpoints from earlier releases may require manually setting the
hybrid layer specification before conversion.

* **Tests**
* Added coverage confirming hybrid specifications survive configuration
serialization and can be recreated successfully.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Keval Morabia <28916987+kevalmorabia97@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
Type of change: Bug fix

`TensorQuantizer.export_amax()` early-returns `self.amax` unsanitized
for dynamic-block
quantizers, while the static path immediately below it has always
substituted `maxbound` for
zero/NaN entries. The `nvfp4` numerics unit sets `type: dynamic`, so a
recipe that applies it to
an *activation* quantizer — e.g.
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`, which targets
`*mlp*input_quantizer` — feeds a raw `0.0` into
`NVFP4QTensor.get_activation_scaling_factor`,
whose assert aborts the entire export:

```
AssertionError: Failed to export module 'model.language_model.layers.37.mlp.gate_proj'
(type=QuantLinear):  activation scaling factor 0.0 not positive.
```

Calibration leaves `amax` at 0 whenever a layer — or an unrouted MoE
expert — saw only zeros, so
one dead layer costs the whole run at the final export step.

This factors the substitution into `_sanitize_export_amax()` and calls
it from both branches. Two
details beyond de-duplication:

- **Branch-free, so it survives a meta `amax`.** `torch.where` +
`nan_to_num` both have meta
kernels; `bool()` on a meta tensor raises. The layerwise and streaming
export flows carry meta
`amax` — `validate_attr` short-circuits on `is_meta` for exactly that
reason — so only the
  warning is gated on a materialized tensor.
- **No longer mutates calibrated state.** The old in-place `amax[amax ==
0] = ...` wrote through a
view of `self._amax`; `torch.where` returns a fresh tensor, so that
hazard disappears.
- **Warns, with a count.** The fix turns a loud failure into a silent
one, and a zero amax means
calibration never activated that layer — worth surfacing rather than
papering over. The message
reports how many entries were substituted, since per-location dedup
otherwise collapses many
dead experts into one uninformative message. A healthy model emits none.

Scope: only the activation path is data-dependent and reachable this
way. Weight-side `_amax` uses
are left alone, since a weight amax of 0 would require an all-zero
weight matrix.

**Knowingly left as follow-up:**
`export/quant_utils.py::get_scaling_factor` discards the sanitized
`amax` when `num_bits == (2, 1)` and recomputes via
`get_weights_scaling_factor_2_from_quantizer`,
which reads `weight_quantizer._amax` raw — so a dynamic-NVFP4 *input*
quantizer on a module whose
*weight* quantizer is a different format (or disabled) can still trip
`assert torch.all(scaling_factor > 0)`. Format dispatch is
weight-driven, so the reported recipe
does not reach that branch; fixing it properly changes a signature
shared with the weight-side
callers and is out of scope here.

Not a regression. The dynamic early return, the `type: dynamic` numerics
unit, and the recipe that
combines them all ship in released 0.46.0 / 0.46.1.

No new or changed API. Exports that previously aborted now complete and
warn:

```python
mtq.quantize(model, quant_cfg, forward_loop)
export_hf_checkpoint(model, export_dir=out)   # before: AssertionError; now: exports + UserWarning
```

- New `test_amax_export_unusable_amax`, parametrized over zero and NaN,
covering the
dynamic-NVFP4 and static per-tensor configs; asserts the exported scale
is positive and that
export leaves the calibrated `amax` untouched. Plus
`test_amax_export_meta_amax`, pinning that
a meta `amax` survives export rather than raising. Both run on CPU and
CUDA via the shared tester.
- `tests/unit/torch/quantization/test_tensor_quantizer_cpu.py` — 40
passed.
`tests/gpu/torch/quantization/test_tensor_quantizer_cuda.py` — 40 passed
(GB300).
- End-to-end repro on GB300, small Llama with one MLP fed all-zero
activations under
`general/ptq/nvfp4_mlp_only-kv_fp8_cast`: dead layer `export_amax()`
`0.0` → `6.0`, live layer
unchanged at `3.921875`, and `export_hf_checkpoint` goes from the
`AssertionError` above to
  writing `model.safetensors`.
- Full `examples/hf_ptq/hf_ptq.py` with the reported recipe and flags on
a healthy model
(Qwen3-0.6B): exits 0 and writes the checkpoint, confirming the normal
path is unaffected.
- `pre-commit run` clean on all changed files.

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
✅
- Did you get Claude approval on this PR?: ✅ — `/claude review` run; its
one IMPORTANT finding (meta-tensor regression) and both SUGGESTIONs
addressed or answered in 251f2e3d

Fixes NVBug 6768300, reported against 0.47.0rc1 on GB200. The reporter
also notes it passed on
0.47.0rc0; that is not explained by code — `git diff
0.47.0rc0..0.47.0rc1` touches
`export/quant_utils.py` only in `get_kv_cache_scaling_factor` (new
`clamp_fp8_scales` argument
whose default preserves the old behaviour) and the INT4-AWQ packing
path, neither of which is on
the dense-HF NVFP4 activation-scale path. Whether `amax` lands on
exactly 0 is
calibration/model-state dependent, which is what makes it look
version-flaky.

Worth flagging separately: in the reported log the **pre-PTQ** sample
output is already
gibberish, so that BF16 checkpoint looks broken independently of
quantization. This change stops
the crash, but such a run will now export a valid-but-garbage checkpoint
— the new warning is the
signal to investigate.

Suggest the `cherry-pick-0.47.0` label so this lands in the ongoing
release.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

* **Bug Fixes**
* Fixed Hugging Face checkpoint export when dynamic-block quantizers
have zero or invalid calibration scales.
* Exports now use a positive fallback scale and issue a warning instead
of failing when applicable.
* Export operations no longer modify the original calibrated quantizer
state.
* Meta-device exports remain non-erroring and preserve device placement.
* **Tests**
* Added coverage for zero- and invalid-scale exports across dynamic and
static quantization modes.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Yue <yueshen@nvidia.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
### What does this PR do?

Type of change: Bug fix: 6701777

Regression source: "[OMNIML-3349] Add FP8 MHA
quantization support for HuggingFace ViT" (#1289), merged into
0.44.0rc3 via the batch cherry-pick #1350. This PR:
  1. Registers nn.LayerNorm as a QuantModule for the first time
     (modelopt/torch/quantization/nn/modules/quant_layernorm.py),
     intended to let FP8_DEFAULT_CFG's BMM input / LayerNorm output
     quantizer rules apply to ViT.
  2. Removes the prior forced Cast-alignment logic in export_onnx.py
that used to normalize Q/DQ node dtypes to trt_high_precision_dtype.

Root Cause:
Once nn.LayerNorm became a registered QuantModule, these wildcards
started unintentionally matching norm1.norm inside FLUX's
AdaLayerNormZero block — an elementwise_affine=False LayerNorm with
no learnable weight/bias. Its input got routed through NVFP4 Q/DQ
(emitted as Float32) while its synthesized affine scale remained
native BFloat16, producing the dtype mismatch.

Chosen fix:
Explicitly exclude nn.LayerNorm from the diffusers NVFP4 presets
rather than touching the global QuantModuleRegistry (which ViT FP8
MHA still needs). Add, in both
modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml and
nvfp4_fp8_mha.yaml, after the existing weight/input wildcard rules
(list order matters — later entries override earlier ones):

  - parent_class: 'nn.LayerNorm'
    quantizer_name: '*'
    enable: false

### Usage

```
python examples/diffusers/quantization/quantize.py --model flux-dev --format fp4 --batch-size 2 --percentile 1.0 --alpha 0.8 --quant-algo max --n-steps 20 --quantized-torch-ckpt-save-path /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4.pt --onnx-dir /tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4 --collect-method default --calib-size 128 --model-dtype BFloat16 --trt-high-precision-dtype BFloat16

trtexec --onnx=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.onnx --builderOptimizationLevel=4 --saveEngine=/tmp/pytest-of-root/pytest-0/test_diffusers_quant_export_on0/flux-dev-fp4/model.plan --stronglyTyped --minShapes=hidden_states:1x1024x64,img_ids:1024x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --optShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1 --maxShapes=hidden_states:1x4096x64,img_ids:4096x3,encoder_hidden_states:1x512x4096,txt_ids:512x3,timestep:1,pooled_projections:1x768,guidance:1
```

### Testing
 The above test commands.

### Before your PR is "*Ready for review*"

Make sure you read and follow [Contributor
guidelines](https://github.com/NVIDIA/Model-Optimizer/blob/main/CONTRIBUTING.md)
and your commits are signed (`git commit -s -S`).

Make sure you read and follow the [Security Best
Practices](https://github.com/NVIDIA/Model-Optimizer/blob/main/SECURITY.md#security-coding-practices-for-contributors)
(e.g. avoiding hardcoded `trust_remote_code=True`, `torch.load(...,
weights_only=False)`, `pickle`, etc.).

- Is this change backward compatible?: N/A
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?:  N/A
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information
N/A

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved Diffusers NVFP4 and NVFP4/FP8 MHA quantization presets by
excluding LayerNorm modules from quantization.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
… disabled quantization during calibration (#2434)

### What does this PR do?

Type of change: Bug fix

During max calibration, `enable_stats_collection()` calls
`disable_quant()`, which sets `_if_quant=False`. The fused causal P-QDQ
attention paths bypass `TensorQuantizer.forward()` and previously
selected the Triton/Kitchen path from the configured enabled state
alone, so P quant-dequant could still execute while quantization was
inactive.

This change:

- enters the Triton P-QDQ path only when `p_bmm_quantizer._if_quant` is
true;
- bypasses all fused P-QDQ paths when the P quantizer is disabled or
quantization is inactive;
- adds focused parameterized regression tests covering Triton and
Kitchen dispatch across enabled, `disable()`, and `disable_quant()`
states.

### Usage

N/A. This restores the existing `disable_quant()` contract and does not
introduce a new API.

### Testing

- `pytest
tests/unit/torch/quantization/plugins/test_attention_quant.py`: 14
passed
- Targeted pre-commit checks passed
- Kitchen coverage verifies the full `{disable(), disable_quant()} ×
{lazy, already initialized}` dispatch matrix remains bypassed while
inactive and resumes after re-enabling
- Reproduced with the exact same ModelOpt 0.47.0rc1 wheel on both sides
on B300: [module regression build
#121](http://dlswqa-nas.nvidia.com:18880/view/yiguo/job/modelopt-quant-module/121/)
- Controlled B300 isolation passed when only the P-BMM quantizers were
hard-disabled during calibration, and also passed when the existing
single Triton attention configuration was forced
- Fixed-code B300 validation passed in three independent runs: [Jenkins
build
123](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/123/),
[Jenkins build
124](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/124/),
and [Jenkins build
125](http://dlswqa-nas.nvidia.com:18880/job/modelopt-quant-module/125/).
Jenkins build 123 compared all 178 module outputs byte-identically and
found 0/377 `amax` and 0/377 `scale` changes.

The confirmed impact is incorrect calibration behavior plus unstable
quantizer state and module outputs; downstream benchmark accuracy impact
has not been established.

### Before your PR is "*Ready for review*"

- Is this change backward compatible?: ✅
- If you copied code from any other sources or added a new PIP
dependency, did you follow guidance in `CONTRIBUTING.md`: N/A
- Did you write any new necessary tests?: ✅
- Did you update
[Changelog](https://github.com/NVIDIA/Model-Optimizer/blob/main/CHANGELOG.rst)?:
N/A
- Did you get Claude approval on this PR?: N/A

### Additional Information

The failure was isolated to fused P-QDQ runtime-state dispatch.
Quantizer topology and configuration were identical in the failing
same-wheel comparison.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **Bug Fixes**
- Corrected attention dispatch so disabled or inactive quantization uses
the original attention implementation.
- Preserved optimized quantized attention when quantization is enabled.
- Ensured attention masks remain unchanged when using the original
attention implementation.
- Improved fallback behavior for fused attention paths, including
correct initialization and reuse when quantization is re-enabled.

- **Tests**
- Added coverage for enabled and disabled quantization states,
attention-mask handling, fallback selection, and result consistency.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Signed-off-by: yingguo-trt <244492186+yingguo-trt@users.noreply.github.com>
Signed-off-by: Chad Voegele <cvoegele@nvidia.com>
@chadvoegele
chadvoegele requested review from a team as code owners September 22, 2026 17:49
@chadvoegele
chadvoegele requested review from a team as code owners September 22, 2026 17:49
@chadvoegele
chadvoegele requested review from ajrasane, h-guo18, kevalmorabia97 and meenchen and removed request for a team September 22, 2026 17:49
@coderabbitai

coderabbitai Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Currently processing new changes in this PR. This may take a few minutes, please wait...

⚙️ Run configuration

Configuration used: Repository: NVIDIA/Model-Optimizer/.coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 44fe0836-45c7-41b2-a521-9325df58c9c9

📥 Commits

Reviewing files that changed from the base of the PR and between 8c3bb9c and eb22ef0.

📒 Files selected for processing (91)
  • CHANGELOG.rst
  • docs/source/deployment/1_tensorrt_llm.rst
  • docs/source/guides/3_pruning.rst
  • docs/source/guides/_customized_model_quantization.rst
  • docs/source/guides/_onnx_quantization.rst
  • examples/diffusers/quantization/ONNX-TRT-Deployment.md
  • examples/diffusers/quantization/config.py
  • examples/diffusers/quantization/onnx_utils/export.py
  • examples/diffusers/quantization/onnx_utils/fp8_onnx_graphsurgeon.py
  • examples/diffusers/quantization/quantize.py
  • examples/diffusers/quantization/utils.py
  • examples/hf_ptq/hf_ptq.py
  • examples/llm_sparsity/weight_sparsity/README.md
  • examples/llm_sparsity/weight_sparsity/export_hf_ckpt.py
  • examples/onnx_ptq/README.md
  • examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py
  • examples/speculative_decoding/main.py
  • examples/speculative_decoding/scripts/ar_validate.py
  • examples/speculative_decoding/scripts/export_hf_checkpoint.py
  • examples/speculative_decoding/scripts/merge_lora.py
  • examples/vllm_serve/Dockerfile
  • examples/vllm_serve/README.md
  • examples/vllm_serve/vllm_ptq_utils.py
  • modelopt/onnx/autocast/convert.py
  • modelopt/onnx/autocast/graphsanitizer.py
  • modelopt/onnx/autocast/precisionconverter.py
  • modelopt/onnx/autocast/referencerunner.py
  • modelopt/onnx/export/fp8_exporter.py
  • modelopt/onnx/export/nvfp4_exporter.py
  • modelopt/onnx/quantization/gs_patching.py
  • modelopt/onnx/quantization/ort_patching.py
  • modelopt/onnx/quantization/quantize.py
  • modelopt/onnx/utils.py
  • modelopt/torch/_deploy/utils/onnx_optimizer.py
  • modelopt/torch/_deploy/utils/torch_onnx.py
  • modelopt/torch/export/hf_export_handlers.py
  • modelopt/torch/nas/plugins/megatron.py
  • modelopt/torch/opt/plugins/mcore_dist_checkpointing.py
  • modelopt/torch/quantization/algorithms.py
  • modelopt/torch/quantization/export_onnx.py
  • modelopt/torch/quantization/model_calib.py
  • modelopt/torch/quantization/model_quant.py
  • modelopt/torch/quantization/nn/modules/quant_module.py
  • modelopt/torch/quantization/nn/modules/tensor_quantizer.py
  • modelopt/torch/quantization/plugins/diffusion/diffusers.py
  • modelopt/torch/quantization/plugins/huggingface.py
  • modelopt/torch/quantization/plugins/megatron.py
  • modelopt/torch/quantization/utils/core_utils.py
  • modelopt/torch/speculative/plugins/hf_training_args.py
  • modelopt/torch/speculative/utils.py
  • modelopt/torch/utils/plugins/__init__.py
  • modelopt/torch/utils/plugins/mbridge.py
  • modelopt/torch/utils/plugins/megatron_layer_specs.py
  • modelopt_recipes/configs/ptq/presets/diffusers/nvfp4.yaml
  • modelopt_recipes/configs/ptq/presets/diffusers/nvfp4_fp8_conv.yaml
  • modelopt_recipes/configs/ptq/presets/diffusers/nvfp4_fp8_mha.yaml
  • modelopt_recipes/general/speculative_decoding/dflash.yaml
  • modelopt_recipes/huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast.yaml
  • modelopt_recipes/huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8.yaml
  • modelopt_recipes/ptq.md
  • plugins/modelopt/skills/deployment/references/trtllm.md
  • tests/_test_utils/torch/megatron/models.py
  • tests/_test_utils/torch/quantization/tensor_quantizer_common.py
  • tests/examples/diffusers/test_diffusers.py
  • tests/examples/hf_ptq/test_hf_ptq_args.py
  • tests/examples/speculative_decoding/test_vllm_hidden_states_aux_layers.py
  • tests/examples/test_onnx_ptq.sh
  • tests/gpu/onnx/test_ort_patching.py
  • tests/gpu/torch/export/test_fsdp2_export.py
  • tests/gpu/torch/quantization/test_onnx_export_cuda.py
  • tests/gpu_megatron/torch/quantization/plugins/test_megatron.py
  • tests/gpu_megatron/torch/utils/plugins/test_mbridge.py
  • tests/gpu_vllm/torch/quantization/test_vllm_dynamic_modules.py
  • tests/unit/examples/test_diffusers_fp4.py
  • tests/unit/onnx/autocast/test_autocast.py
  • tests/unit/onnx/autocast/test_graphsanitizer.py
  • tests/unit/onnx/autocast/test_precisionconverter.py
  • tests/unit/onnx/autocast/test_referencerunner.py
  • tests/unit/onnx/quantization/test_ort_patching_histogram.py
  • tests/unit/onnx/quantization/test_qdq_utils.py
  • tests/unit/onnx/test_autocast_quantize.py
  • tests/unit/onnx/test_onnx_utils.py
  • tests/unit/recipe/test_step3p7_recipes.py
  • tests/unit/torch/deploy/utils/test_torch_onnx_utils.py
  • tests/unit/torch/quantization/plugins/test_attention_quant.py
  • tests/unit/torch/quantization/plugins/test_fused_experts.py
  • tests/unit/torch/quantization/plugins/test_moe_linear.py
  • tests/unit/torch/quantization/test_autoquant.py
  • tests/unit/torch/quantization/test_onnx_export_cpu.py
  • tests/unit/torch/quantization/test_quantize_cpu.py
  • tools/launcher/examples/nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16/offline_kd_qad.yaml
 ______________________________________________________________
< Your function has more side effects than my caffeine intake. >
 --------------------------------------------------------------
  \
   \   (\__/)
       (•ㅅ•)
       /   づ
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-09-22 17:55 UTC

@codecov

codecov Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.17557% with 41 lines in your changes missing coverage. Please review.
✅ Project coverage is 78.64%. Comparing base (8c3bb9c) to head (eb22ef0).

Files with missing lines Patch % Lines
modelopt/torch/speculative/utils.py 44.44% 25 Missing ⚠️
modelopt/onnx/autocast/referencerunner.py 87.87% 4 Missing ⚠️
modelopt/torch/quantization/plugins/huggingface.py 92.15% 4 Missing ⚠️
modelopt/onnx/utils.py 97.33% 2 Missing ⚠️
modelopt/torch/quantization/plugins/megatron.py 84.61% 2 Missing ⚠️
modelopt/torch/quantization/utils/core_utils.py 66.66% 2 Missing ⚠️
...lopt/torch/opt/plugins/mcore_dist_checkpointing.py 75.00% 1 Missing ⚠️
modelopt/torch/quantization/algorithms.py 99.30% 1 Missing ⚠️
Additional details and impacted files
@@                Coverage Diff                 @@
##           release/0.47.0    #2502      +/-   ##
==================================================
- Coverage           78.68%   78.64%   -0.05%     
==================================================
  Files                 526      527       +1     
  Lines               61331    61635     +304     
==================================================
+ Hits                48259    48473     +214     
- Misses              13072    13162      +90     
Flag Coverage Δ
examples-diffusers 21.10% <17.74%> (+0.39%) ⬆️
examples-gpt-oss 13.19% <7.06%> (-0.08%) ⬇️
examples-hf_ptq 21.39% <34.35%> (-0.14%) ⬇️
examples-llm_distill 13.25% <7.06%> (-0.08%) ⬇️
examples-llm_eval 17.00% <13.35%> (-0.09%) ⬇️
examples-llm_sparsity 15.79% <7.06%> (-0.12%) ⬇️
examples-specdec_bench 12.93% <7.06%> (-0.07%) ⬇️
examples-speculative_decoding 17.45% <16.22%> (-0.14%) ⬇️
examples-torch_onnx 21.79% <49.61%> (-0.01%) ⬇️
examples-torch_trt 14.99% <10.11%> (-0.08%) ⬇️
gpu 58.38% <62.02%> (-0.66%) ⬇️
regression 14.84% <10.68%> (+0.01%) ⬆️
unit 56.70% <85.87%> (+0.60%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.