MiniMax H3 FL2VA Pruned INT8 ConvRot + TaoMate 8-Step

A custom merged MiniMax H3 FL2VA model combining the official pruned INT8 ConvRot-256 model with the TaoMate-H3 3-step LoRA.

The TaoMate adaptation is permanently merged into the model weights, so the separate TaoMate LoRA is not required during inference.

This release is specifically optimized and tested for low-step audio-video generation in ComfyUI.

Model Overview

Model file

minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors

Base model

Comfy-Org/MiniMax-H3

Architecture

MiniMax H3 FL2VA

Format

SafeTensors

Main transformer quantization

  • INT8
  • ConvRot
  • Group size: 256
  • Row-wise scaling

Higher-precision components

  • AdaLN projections: FP16
  • Token Refiner: BF16
  • Token Refiner was not INT8-quantized

TaoMate Integration

This model was created by merging the TaoMate-H3 3-step LoRA directly into the official pruned MiniMax H3 FL2VA INT8 ConvRot-256 weights.

Official TaoMate-H3 project:

https://github.com/TaoLiveAIGC/TaoMate-H3

The original TaoMate adapter uses:

  • LoRA rank: 128
  • LoRA alpha: 128
  • Native alpha/rank scale: 1.0
  • Merge strength used for this release: 0.9

The LoRA contains:

  • 416 LoRA tensors
  • 208 target modules

Target distribution:

  • 200 main transformer targets
  • 8 Token Refiner targets

The merge includes the Token Refiner LoRA targets as well as the main transformer targets.


Merge Method

The official pruned INT8 ConvRot-256 MiniMax H3 FL2VA model was used as the starting point.

The TaoMate LoRA delta was calculated from the low-rank matrices and merged into the corresponding base weights.

For the main INT8 ConvRot layers, the existing quantized representation was reconstructed, the LoRA delta was applied in BF16 precision, and the resulting weights were requantized to INT8 ConvRot-256 using row-wise scaling.

The Token Refiner remained in BF16 and its eight TaoMate target modules were merged directly into the BF16 weights.

AdaLN parameters remained in FP16/high precision.

The final result is a single approximately 19.5 GB SafeTensors model with the TaoMate adaptation already integrated.


Why This Model Exists

The purpose of this release is to provide a practical single-file MiniMax H3 FL2VA model for ComfyUI that combines:

  • MiniMax H3 audio-video generation
  • TaoMate low-step adaptation
  • INT8 ConvRot-256 quantization
  • Integrated TaoMate LoRA
  • No separate LoRA loading
  • Low-step inference
  • Audio and video generation in one model

This is particularly useful for ComfyUI users who want TaoMate-style low-step generation without loading the large TaoMate LoRA separately for every generation.

Important: This merged model contains the TaoMate LoRA weights, but it does not reproduce every component of the official TaoMate streaming runtime. The official TaoMate runtime contains additional streaming and KV-cache scheduling logic.


Recommended Settings

After testing multiple configurations, the recommended configuration for this merged model is:

Parameter Recommended
Steps 8
Sampler Euler
Scheduler Simple
CFG 1.0
Video Sigma Shift 8
Audio Sigma Shift 3
Denoise 1.0
Separate TaoMate LoRA No

Recommended Preset

Steps:          8
Sampler:        Euler
Scheduler:      Simple
CFG:            1.0

Video Sigma:    8
Audio Sigma:    3

Denoise:        1.0

The TaoMate LoRA is already merged into this model.

Do not load the original TaoMate LoRA again.


Why 8 Steps?

The original TaoMate adaptation is a 3-step distillation.

However, testing this merged ComfyUI model showed that 3 sampling steps were not optimal for visual quality.

Observed behavior:

3 steps
β†’ visible instability in facial details, eyes, motion and fine structures

6 steps
β†’ significantly improved, but some residual defects remained

8 steps
β†’ noticeably more complete detail reconstruction and better structural stability

Therefore this release is specifically optimized for 8 sampling steps.

The underlying adaptation remains the TaoMate 3-step LoRA.


Sigma Shift Testing

Three test videos were generated to compare different Video Sigma Shift values.

The Audio Sigma Shift was kept at 3.

Video Sigma Shift 12 / Audio 3

Watch Sigma Shift 12/3

Video Sigma Shift 6 / Audio 3

Watch Sigma Shift 6/3

Video Sigma Shift 8 / Audio 3 β€” Recommended

Watch Sigma Shift 8/3

The 8/3 configuration provided the best observed balance between:

  • facial stability
  • eye detail
  • fine structures
  • motion consistency
  • overall visual quality

Therefore:

Video Sigma Shift = 8

Audio Sigma Shift = 3


Benchmark

Benchmark conditions:

  • Video duration: 5 seconds
  • Resolution: approximately 1 MP
  • Sampling steps: 8
  • ComfyUI generation
  • Same local hardware/environment for both tests

This merged model

minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors

Generation time:

225 seconds

Reference configuration

minimax_h3_fl2va_pruned_int8_convrot.safetensors

minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors

Generation time:

239 seconds

Difference

225 seconds vs 239 seconds

Difference: 14 seconds

Relative difference: approximately 5.9%

The merged model completed this particular benchmark approximately 5.9% faster than the tested INT8 + separate BF16 8-step reference configuration.

This is a local benchmark and should not be considered a universal performance guarantee. Performance can vary depending on GPU, VRAM, CUDA, PyTorch, ComfyUI version, resolution, workflow and other runtime conditions.


Benchmark Comparison

Watch benchmark comparison

The comparison shows this merged model against:

minimax_h3_fl2va_pruned_int8_convrot.safetensors

minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors


Example Videos

Sigma Shift comparisons

Performance comparison


Quantization

Main transformer:

Quantization:       INT8
Format:             ConvRot
Group size:         256
Scaling:            Row-wise

The model preserves higher precision for components where maintaining numerical behavior is important:

AdaLN:              FP16
Token Refiner:      BF16

The Token Refiner is intentionally not represented as INT8 ConvRot in this release.


Model Structure

MiniMax H3 FL2VA
β”‚
β”œβ”€β”€ Main Transformer
β”‚   β”œβ”€β”€ 50 transformer blocks
β”‚   β”œβ”€β”€ INT8 ConvRot-256
β”‚   └── TaoMate merged into 200 target modules
β”‚
β”œβ”€β”€ Token Refiner
β”‚   β”œβ”€β”€ 2 blocks
β”‚   β”œβ”€β”€ BF16
β”‚   └── TaoMate merged into 8 target modules
β”‚
β”œβ”€β”€ AdaLN
β”‚   └── FP16
β”‚
└── Other H3 conditioning / patch components

Total TaoMate integration:

LoRA tensors:       416
Target modules:     208
Main targets:       200
Token Refiner:      8
Rank:               128
Alpha:              128
Merge strength:     0.9

Recommended ComfyUI Configuration

Model:
minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors

Steps:
8

Sampler:
Euler

Scheduler:
Simple

CFG:
1.0

Video Sigma Shift:
8

Audio Sigma Shift:
3

Denoise:
1.0

No additional TaoMate LoRA is required.


Base Model

Comfy-Org MiniMax H3:

https://huggingface.co/Comfy-Org/MiniMax-H3

Official MiniMax H3:

https://huggingface.co/MiniMaxAI/MiniMax-H3

TaoMate-H3:

https://github.com/TaoLiveAIGC/TaoMate-H3


Checksums

Original Comfy-Org pruned INT8 ConvRot base:

SHA256:
e889202c41dafb67b10d67b97f0d8541508036a6090af23425a5c2615d03c47a

This merged release:

SHA256:
302041433520d0ae7386fbe380cbd439d296b7ff9b1217537a37921e72c1e407

License

This model is a community derivative/merged release based on MiniMax H3 and TaoMate-H3 components.

Users must review and comply with the applicable licenses of the base model and the TaoMate-H3 project before redistribution or commercial use.

The Comfy-Org MiniMax H3 distribution identifies the MiniMax H3 Community License Agreement as its applicable license.


Disclaimer

This is an experimental community merge optimized and tested for ComfyUI.

The recommended 8-step / 8-3 configuration is based on empirical testing of this specific merged release.

Results may vary depending on hardware, software versions, workflow, resolution, generation length and other runtime parameters.

Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Asirus/minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step

Finetuned
(20)
this model