Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
28 commits
Select commit Hold shift + click to select a range
613175a
Speed up FSDP2 MoE calibration by dropping redundant expert-weight ga…
sugunav14 Sep 9, 2026
5d4d395
feat(quantization): PTQ support for Step-3.7 MoE checkpoints (#2202)
Edwardf0t1 Sep 10, 2026
839fba6
[6410139] Fix ONNX AutoCast for large external initializers (#2317)
ajrasane Sep 10, 2026
7800218
specdec: config_overrides for nested text_config checkpoints + load V…
yeyu-nvidia Sep 10, 2026
d19c7d6
[6508436] Fix BF16 FP8 ONNX export (#2314)
ajrasane Sep 10, 2026
2fbcca2
deprecate trtllm-build in weight_sparsity (#2371)
noeyy-mino Sep 11, 2026
68db2c9
Fix protobuf size-check failures in ONNX deployment (#2403)
noeyy-mino Sep 11, 2026
aca3f13
Document the MSE calibration API (#2405)
realAsma Sep 11, 2026
c51568c
Fix TEGroupedMLP quantizer checkpoint resharding (#2319)
jenchen13 Sep 11, 2026
e75d6c1
feat(quantization): fail fast when a quant config matches no weight q…
Edwardf0t1 Sep 11, 2026
83876db
[6701308][OMNIML-5805] Correct ONNX PTQ documentation contracts (#2413)
ajrasane Sep 11, 2026
5ecede0
[6463897] Fix narrow FP16 histogram calibration (#2412)
ajrasane Sep 12, 2026
72a53b0
Fix distributed AutoQuantize scoring and share backward setup (#2231)
joshua-hill Sep 13, 2026
60638ec
Document the nvfp4_act_headroom calibration variant in ptq.md (#2439)
cjluo-nv Sep 16, 2026
9880ff9
docs: replace legacy TensorRT-LLM engine deployment guidance (#2436)
chadvoegele Sep 16, 2026
6ea33c5
Fail fast on non-finite AutoQuantize output gradients (#2432)
meenchen Sep 17, 2026
763cf95
fix(specdec): resolve the eagle aux-layer preset in the vLLM hidden-s…
yeyu-nvidia Sep 17, 2026
3d9872f
Fix vLLM fakequant calibration for hybrid attention models (#2414)
kinjalpatel27 Sep 17, 2026
961741c
docs(recipes): point chat_template at where the templates actually li…
yeyu-nvidia Sep 17, 2026
0470b12
[6771663] Preserve ONNX API output types when wiring casts (#2451)
ajrasane Sep 18, 2026
a2cee6a
[5565357] Fix SDXL NVFP4 export and performance (#2336)
ajrasane Sep 18, 2026
a48fd62
Move Nemotron Nano offline KD example (#2471)
jenchen13 Sep 18, 2026
415e826
docs: clarify canonical pruning documentation source (#1871) (#2469)
DIYA73 Sep 18, 2026
39a61af
Fix hf_ptq.py discarding completed PTQ run on sanity-generate() failu…
Edwardf0t1 Sep 19, 2026
b935126
Fix hybrid stack spec serialization in Megatron-Bridge checkpoints (#…
kevalmorabia97 Sep 19, 2026
042e1ff
Fix HF export crash when a dynamic-block quantizer has zero amax (#2438)
yueshen2016 Sep 21, 2026
6412746
Noeyy/fix bug 6701777 (#2402)
noeyy-mino Sep 22, 2026
eb22ef0
[https://nvbugspro.nvidia.com/bug/6778095] Fix fused P-QDQ to respect…
yingguo-trt Sep 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions CHANGELOG.rst
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,8 @@ Changelog
- Add ``mtq.temporarily_fold_weights`` for repeated frozen-weight inference and ``mtq.preserve_quantizer_attributes_context`` for restoring temporary quantizer property and type changes. Temporary folding snapshots affected fake-quant weights on a configurable device and restores them with their quantizer state; retained pre-quant scales are inactive, while shared weights, shared quantizers, and ``SequentialQuantizer`` weights are unsupported.
- Add the ``nvfp4_act_headroom`` calibration algorithm for NVFP4 **activation** global scales. Instead of setting the global scale from the largest per-block amax seen during calibration (plain ``max``, which leaves no room above it so any larger activation saturates), it anchors the scale to a low percentile of the per-block amax distribution, leaving the rest of the FP8 block-scale range as headroom: ``amax = max(rho * anchor, upper)``, where ``anchor`` and ``upper`` are the per-block amaxes at ``anchor_percentile`` (default 1) and ``upper_percentile`` (default 99.99; set to 100 to never clip calibration data), and ``rho`` (default 16384) is the headroom factor. Applies only to NVFP4 dynamic-block input quantizers; ``SequentialQuantizer`` activation quantizers raise. Weight scales are an orthogonal axis selected by a nested ``weight_scale_algorithm`` (``max`` by default, or ``mse`` / ``local_hessian``), so one recipe can combine a weight calibration with this activation policy in a single pass. Ships ``modelopt_recipes/general/ptq/nvfp4_act_headroom-kv_fp8_cast.yaml``, which mirrors ``nvfp4_default-kv_fp8_cast`` with only the calibration algorithm swapped and exports a standard NVFP4 checkpoint.

- Add PTQ support for Step-3.7 (``stepfun-ai/Step-3.7-Flash``), whose routed experts were previously left unquantized. Quantize with the new ``huggingface/step3p7/ptq/nvfp4_experts_only-kv_fp8_cast`` or ``huggingface/step3p7/ptq/nvfp4_mlp_only-kv_fp8`` recipes rather than the general ones, which select experts by module names Step does not use.

*Megatron Framework (M-LM / M-Bridge)*

- Add ``clamp_kv_cache_scales`` to ``export_mcore_gpt_to_hf``. Set it to ``False`` when exporting a QAT Megatron-Core model to preserve its learned FP8 KV-cache scales; the default retains the existing minimum scale of 1.0.
Expand All @@ -34,6 +36,7 @@ Changelog

**Backward Breaking Changes**

- ``get_te_hybrid_stack_spec`` was removed from ``modelopt.torch.nas.plugins.megatron``; it had no use outside tests. Use ``modelopt.torch.utils.plugins.megatron_layer_specs.te_hybrid_stack_spec_sequential_mlp`` for the SequentialMLP layout, or ``megatron.core.models.hybrid.hybrid_layer_specs.hybrid_stack_spec`` for grouped GEMM.
- Migrate the FAR3D ONNX PTQ example to the shared evaluator and ModelOpt containers and ``quantize_vovnet.py``. Only the encoder supports INT8 and FP8; decoder calibration, quantization, and related CLI flags are removed, and the decoder remains in its exported mixed FP16/FP32 precision.
- Image-text calibration with ``--calib_with_images`` now forwards multimodal batches through the complete VLM for all VLM families, so existing non-Nemotron commands may produce different language-model activation ranges and output scales. Recipe-based VLM PTQ also targets the complete VLM: vision modules stay in high precision by default and are quantized only when a model-specific recipe enables them, so custom recipes must explicitly exclude vision modules when required.
- Move the checkpoint-mirror recipe tier from ``huggingface/models/<org>/<checkpoint>/`` to the top-level ``models/<org>/<model_id>/``, keyed by each recipe's canonical Hugging Face Hub id — so the Step 3.5 Flash recipe moves to ``models/stepfun-ai/Step-3.5-Flash/ptq/`` and the NVIDIA Nemotron recipes gain the ``NVIDIA-`` prefix (e.g. ``models/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16/ptq/nvfp4-mse``). Update any saved ``--recipe`` paths for these checkpoint recipes accordingly; the per-``model_type`` recipes under ``huggingface/`` are unchanged.
Expand All @@ -45,6 +48,7 @@ Changelog
- Remove in-trainer quantization via ``QuantizationArguments.quant_cfg`` / ``--quant_cfg`` (deprecated in 0.45); use ``--recipe``. New recipes ``general/ptq/mxfp4_mlp_weight_only`` and ``general/ptq/nvfp4_mlp_weight_only`` replace ``MXFP4_MLP_WEIGHT_ONLY_CFG`` / ``NVFP4_MLP_WEIGHT_ONLY_CFG`` in the ``examples/gpt-oss`` QAT flow.
- Remove the ``QuantizationArgumentsWithConfig`` alias in ``modelopt.torch.quantization.plugins.transformers_trainer`` (deprecated in 0.45). Use ``QuantizationArguments``.
- Transformer Engine ``TEGroupedLinear`` (fused MoE experts) now uses **per-expert** weight quantization (one ``amax`` per expert) instead of a single shared ``amax``, so ModelOpt checkpoints containing quantized ``TEGroupedLinear`` modules saved before 0.47 are **not compatible** with 0.47. Re-run PTQ to regenerate compatible checkpoints.
- ``mtq.quantize`` now raises when a config asks for weight quantization but none of its weight-quantizer patterns match the model, instead of calibrating and exporting a silently unquantized checkpoint (``"quant_algo": null``). Configs that quantize activations or the KV cache only are unaffected, as are patterns that match and are then disabled by a later entry. If this fires, use the recipe for that architecture under ``modelopt_recipes/huggingface/<model_type>/`` or fix the module patterns. Set ``MODELOPT_SKIP_WEIGHT_QUANT_CHECK=1`` to disable the check process-wide, e.g. for a pipeline-parallel rank whose local stage legitimately has none of the targeted modules.

**Deprecations**

Expand All @@ -53,8 +57,18 @@ Changelog

**Bug Fixes**

- Fix shared ONNX export metadata and Diffusers attention policy: every ``NVFP4QuantExporter`` post-process now upgrades the default-domain opset to at least 23, all FP8 custom-op exports re-run ONNX shape/type inference after setting output metadata, and quantized SDPA derives FP8 MHA enablement from the live Q/K/V quantizers instead of honoring a caller-set ``_disable_fp8_mha`` attribute.
- Fix ONNX FP16 conversion failing to preserve public output types when type inference changes a graph output declaration before output casts are inserted.
- Fix ``examples/hf_ptq/hf_ptq.py`` discarding a completed PTQ run (no checkpoint exported) when the optional post-quantization sanity-check ``generate()`` call raised, for example because ``device_map="auto"`` placed part of the model on CPU. That failure is now caught and only skips the sanity check; export proceeds regardless.
- Hybrid (e.g. Nemotron-H) checkpoints saved by the ``examples/megatron_bridge`` scripts now record their layer spec in ``run_config.yaml`` in a form that reloads, so they can be converted to HuggingFace; a checkpoint saved by an earlier release still needs its ``model.hybrid_stack_spec`` block replaced by hand.
- Fail fast on non-finite AutoQuantize output gradients with an actionable error before accumulating sensitivity scores, without changing attention backend settings.
- Fix ONNX INT8 entropy calibration failing or producing invalid quantization parameters for FP16 activations.
- Fix HuggingFace checkpoint export failing with ``activation scaling factor 0.0 not positive`` when a dynamic-block quantizer (such as an NVFP4 input quantizer) ends calibration with an amax of zero because the calibration data never activated that layer or expert. Such a quantizer now exports a positive fallback scale and warns instead of crashing, matching what static quantizers already did; if you see the warning, check whether the layer is expected to be inactive and consider a larger calibration size.
- Speed up ``mtq.quantize`` on FSDP2-sharded fused-MoE models. Promoting static-block weight quantizers gathered each expert's slice of the fused weight across ranks even though only quantizer state is read, adding a collective per expert to calibration.
- Fix ONNX AutoCast failing on models with external initializers larger than 2 GiB.
- Avoid querying CUDA/Blackwell capability when ``NVFP4QTensor.quantize`` uses its CPU path or has the optional TensorRT-LLM fast path disabled.
- Fix NVFP4 ONNX export to quantize FP4 weights with the published FP8 block scales, matching eager ModelOpt packed weights. Block scales below ``2**-9`` are now clamped to that minimum, and non-finite or negative scales raise an error.
- Fix FP8 ONNX export of BF16 models during real-weight compression.
- Fix Megatron-Bridge Quantization Aware Distillation of a vision-language model silently discarding the ModelOpt state, so the distilled checkpoint restored no quantizers and exported as an unquantized model. Re-run QAD to regenerate any affected checkpoint.
- Fix Megatron-Core HuggingFace export silently omitting fused (grouped GEMM) MoE experts for architectures without an ``experts.linear_fc1`` rule (e.g. ``Qwen3MoeForCausalLM``), which produced a valid-looking checkpoint containing no expert weights. The exporter now raises instead of writing that checkpoint; the scripts also avoid the situation by selecting ``SequentialMLP`` for those architectures.
- Fix GatedDeltaNet (Qwen3.5) quantizer exclusions on Megatron-Core: the recipe patterns name the HuggingFace ``linear_attn`` module, so the ``conv1d`` was calibrated and the alpha / beta gate projections were exported in FP8. ``conv1d`` now has a ``self_attention`` alias in the default disabled-quantizer units, and the alpha / beta projections are exported in BF16 (they share Megatron's fused ``in_proj`` quantizer and cannot be disabled by name).
Expand All @@ -71,6 +85,7 @@ Changelog
- Fix EAGLE-3 training with context parallelism (``--cp_size > 1`` in ``examples/speculative_decoding``), which failed to start on ``accelerate >= 1.13`` and then raised ``got mixed torch.Tensor and DTensor``.
- Polygraphy minimum dependency upgraded to ``0.53.4`` to solve ONNX AutoCast failures when marking optional graph outputs.
- Fix ``--kv_cache_free_gpu_memory_fraction`` having no effect on the ``lm_eval`` task of ``examples/hf_ptq/scripts/huggingface_example.sh``, where the KV cache always took TensorRT-LLM's default 90% of free GPU memory and evaluation could run out of memory. ``examples/llm_eval/lm_eval_trtllm.py`` now takes ``kv_cache_free_gpu_memory_fraction`` in ``--model_args``, defaulting to 0.8.
- Fix ``--aux-layers eagle`` failing in the vLLM offline hidden-state dump (``examples/speculative_decoding/collect_hidden_states/compute_hidden_states_vllm.py``). ``eagle`` is the flag's default, but the dump's standalone resolver -- a copy kept so the script runs in a stock vLLM container without ModelOpt -- only handled ``dflash`` and explicit id lists, so the documented invocation aborted with ``invalid literal for int(): 'eagle'`` before any state was written. An unrecognised preset now reports which values are accepted instead of surfacing the raw ``int()`` error.

0.46 (2026-08-17)
^^^^^^^^^^^^^^^^^
Expand Down
159 changes: 13 additions & 146 deletions docs/source/deployment/1_tensorrt_llm.rst
Original file line number Diff line number Diff line change
Expand Up @@ -2,149 +2,16 @@
TensorRT-LLM
==========================

**Deprecation Notice**: The export_tensorrt_llm_checkpoint API will be deprecated in future releases. Users are encouraged to transition to the :doc:`unified HF export API <3_unified_hf>`, which provides enhanced functionality and flexibility for exporting models to multiple inference frameworks including TensorRT-LLM, vLLM, and SGLang.

.. note::

Please read the `TensorRT-LLM checkpoint workflow <https://github.com/NVIDIA/TensorRT-LLM/blob/main/docs/source/legacy/architecture/checkpoint.md>`_
first before going through this section.



ModelOpt toolkit supports automatic conversion of ModelOpt exported LLM to the TensorRT-LLM checkpoint and the engines for accelerated inferencing.

This conversion is achieved by:

#. Converting Huggingface, Megatron-Bridge and ModelOpt exported checkpoints to the TensorRT-LLM checkpoint.
#. Building TensorRT-LLM engine from the TensorRT-LLM checkpoint.


Export Quantized Model
======================

After the model is quantized, the quantized model can be exported to the TensorRT-LLM checkpoint format stored as

#. A single JSON file recording the model structure and metadata (config.json)
#. A group of safetensors files, each recording the local calibrated model on a single GPU rank (model weights, scaling factors per GPU).

The export API (:meth:`export_tensorrt_llm_checkpoint <modelopt.torch.export.model_config_export.export_tensorrt_llm_checkpoint>`) can be used as follows:

.. code-block:: python

from modelopt.torch.export import export_tensorrt_llm_checkpoint

with torch.inference_mode():
export_tensorrt_llm_checkpoint(
model, # The quantized model.
decoder_type, # The type of the model as str, e.g gpt, gptj, llama.
dtype, # the weights data type to export the unquantized layers.
export_dir, # The directory where the exported files will be stored.
inference_tensor_parallel, # The number of GPUs used in the inference time for tensor parallelism.
inference_pipeline_parallel, # The number of GPUs used in the inference time for pipeline parallelism.
)

If the :meth:`export_tensorrt_llm_checkpoint <modelopt.torch.export.model_config_export.export_tensorrt_llm_checkpoint>` call is successful, the TensorRT-LLM checkpoint will be saved. Otherwise, e.g. the ``decoder_type`` is not supported, a torch state_dict checkpoint will be saved instead.

.. list-table:: Model support matrix for the TensorRT-LLM checkpoint export
:header-rows: 1

* - Model / Quantization
- FP16 / BF16
- FP8
- INT8_SQ
- INT4_AWQ
* - GPT2
- Yes
- Yes
- Yes
- No
* - GPTJ
- Yes
- Yes
- Yes
- Yes
* - LLAMA 2
- Yes
- Yes
- Yes
- Yes
* - LLAMA 3
- Yes
- Yes
- No
- Yes
* - Mistral
- Yes
- Yes
- Yes
- Yes
* - Mixtral 8x7B
- Yes
- Yes
- No
- Yes
* - Falcon 40B, 180B
- Yes
- Yes
- Yes
- Yes
* - Falcon 7B
- Yes
- Yes
- Yes
- No
* - MPT 7B, 30B
- Yes
- Yes
- Yes
- Yes
* - Baichuan 1, 2
- Yes
- Yes
- Yes
- Yes
* - ChatGLM2, 3 6B
- Yes
- No
- No
- Yes
* - Bloom
- Yes
- Yes
- Yes
- Yes
* - Phi-1, 2, 3
- Yes
- Yes
- Yes
- Yes
* - Nemotron 8
- Yes
- Yes
- No
- Yes
* - Gemma 2B, 7B
- Yes
- Yes
- No
- Yes
* - Recurrent Gemma
- Yes
- Yes
- Yes
- Yes
* - StarCoder 2
- Yes
- Yes
- Yes
- Yes
* - Qwen-1, 1.5
- Yes
- Yes
- Yes
- Yes

Convert to TensorRT-LLM
=======================

Once the TensorRT-LLM checkpoint is available, please follow the `TensorRT-LLM build API <https://github.com/NVIDIA/TensorRT-LLM/blob/main/docs/source/legacy/architecture/workflow.md#build-apis>`_ to build and deploy the quantized LLM.
For current TensorRT-LLM deployments, export quantized models with
:meth:`export_hf_checkpoint <modelopt.torch.export.unified_export_hf.export_hf_checkpoint>`
and load the exported Hugging Face checkpoint with TensorRT-LLM's PyTorch backend.
See the :doc:`unified HF export guide <3_unified_hf>` for export and deployment
examples, supported models, and quantization formats. This workflow does not require
building a TensorRT engine.

.. warning::

The ``export_tensorrt_llm_checkpoint`` API exports checkpoints for the legacy
TensorRT backend, which current TensorRT-LLM releases no longer support.
The API will be deprecated in a future release.
Use ``export_hf_checkpoint`` instead.
7 changes: 7 additions & 0 deletions docs/source/guides/3_pruning.rst
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,13 @@ Pruning
Checkout `Megatron-Bridge Minitron Pruning & Distillation <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/megatron_bridge>`_ and
`ResNet20 on CIFAR-10 Notebook <https://github.com/NVIDIA/Model-Optimizer/blob/main/examples/pruning/cifar_resnet.ipynb>`_
for an end-to-end example of pruning.
.. note::

For Minitron (LLM/VLM pruning via Megatron-Bridge/Megatron-LM) and Puzzletron pruning,
the canonical reference is the
`Pruning README <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/pruning>`_
which contains up-to-date API examples, support matrix, guidelines, and distillation
hyperparameters. This guide covers FastNAS pruning for Computer Vision models.

ModelOpt provides three main pruning methods (aka ``mode``) - Minitron, Puzzletron, and FastNAS - via a unified API
:meth:`mtp.prune <modelopt.torch.prune.pruning.prune>`. Given a model,
Expand Down
2 changes: 1 addition & 1 deletion docs/source/guides/_customized_model_quantization.rst
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ As ModelOpt cannot detect these linear ops out-of-the-box, a HugggingFace plugin
#. Rewrite the linear ops (w1, v1 and v2) as a standard ``nn.Linear`` op, and re-implement the ``forward`` method.
#. Register the new dynamic ``_QuantDbrxExperts`` to replace the ``DbrxExperts`` from the modeling_dbrx.py in the ``transformers`` library
#. Try quantize the DBRX model after the plugin is implemented, feel free to follow the `hf_ptq example <https://github.com/NVIDIA/Model-Optimizer/tree/main/examples/hf_ptq>`_.
#. TensorRT-LLM is open-sourced. If this customized model is not supported by TensorRT-LLM yet, please modify :meth:`export_tensorrt_llm_checkpoint <modelopt.torch.export.export_tensorrt_llm_checkpoint>` or :meth:`export_hf_checkpoint <modelopt.torch.export.export_hf_checkpoint>` to export the quantized model for deployment with a customized TensorRT-LLM modeling implementation. Feel free to :doc:`contact us <../support/1_contact>` if further support is needed.
#. Export the quantized model with :meth:`export_hf_checkpoint <modelopt.torch.export.unified_export_hf.export_hf_checkpoint>`. If the customized model is not supported by TensorRT-LLM, add support in its PyTorch backend and adapt the HF exporter if needed. See the :doc:`unified HF export guide <../deployment/3_unified_hf>` or :doc:`contact us <../support/1_contact>` for help.

The following code snippet is excerpted from ``modelopt/torch/quantization/plugins/huggingface.py``

Expand Down
2 changes: 1 addition & 1 deletion docs/source/guides/_onnx_quantization.rst
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ Requirements
Apply Post Training Quantization (PTQ)
======================================

PTQ should be done with a calibration dataset. If calibration dataset is not provided, ModelOpt will use random scales for the QDQ nodes.
PTQ should be done with a calibration dataset. Random calibration inputs are used when no calibration dataset is provided.

Prepare calibration dataset
---------------------------
Expand Down
4 changes: 2 additions & 2 deletions examples/diffusers/quantization/ONNX-TRT-Deployment.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,12 +28,12 @@ python quantize.py \

#### FLUX-Dev|SDXL|SDXL-Turbo|LTX-Video FP8/FP4 [Script](./quantize.py)

*In our example code, FP4 is only supported for Flux. However, you can modify our script to enable FP4 format support for your own model.*
FP4 ONNX export is supported for Flux and SDXL.

```sh
python quantize.py \
--model {flux-dev|sdxl-1.0|sdxl-turbo|ltx-video-dev} --model-dtype {Half|BFloat16} --trt-high-precision-dtype {Half|BFloat16} \
--format {fp8|fp4} --batch-size 2 --calib-size {128|256} --quantize-mha \
--format {fp8|fp4} --batch-size 2 --calib-size {128|256} \
--n-steps 20 --quantized-torch-ckpt-save-path ./{MODEL_NAME}.pt --collect-method default \
--onnx-dir {ONNX_DIR}
```
Expand Down
3 changes: 3 additions & 0 deletions examples/diffusers/quantization/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,9 @@
NVFP4_FP8_MHA_CONFIG = load_config(
"configs/ptq/presets/diffusers/nvfp4_fp8_mha", schema_type=QuantizeConfig
).model_dump(exclude_unset=True)
NVFP4_FP8_CONV_CONFIG = load_config(
"configs/ptq/presets/diffusers/nvfp4_fp8_conv", schema_type=QuantizeConfig
).model_dump(exclude_unset=True)


def set_quant_config_attr(quant_config, trt_high_precision_dtype, quant_algo, **kwargs):
Expand Down
Loading
Loading