- MiniMax H3 FL2VA Pruned INT8 ConvRot + TaoMate 8-Step
- TaoMate Integration
- Merge Method
- Why This Model Exists
- Recommended Settings
- Why 8 Steps?
- Sigma Shift Testing
- Benchmark
- Benchmark Comparison
- Example Videos
- Quantization
- Model Structure
- Recommended ComfyUI Configuration
- Base Model
- Checksums
- License
- Disclaimer
MiniMax H3 FL2VA Pruned INT8 ConvRot + TaoMate 8-Step
A custom merged MiniMax H3 FL2VA model combining the official pruned INT8 ConvRot-256 model with the TaoMate-H3 3-step LoRA.
The TaoMate adaptation is permanently merged into the model weights, so the separate TaoMate LoRA is not required during inference.
This release is specifically optimized and tested for low-step audio-video generation in ComfyUI.
Model Overview
Model file
minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors
Base model
Comfy-Org/MiniMax-H3
Architecture
MiniMax H3 FL2VA
Format
SafeTensors
Main transformer quantization
- INT8
- ConvRot
- Group size: 256
- Row-wise scaling
Higher-precision components
- AdaLN projections: FP16
- Token Refiner: BF16
- Token Refiner was not INT8-quantized
TaoMate Integration
This model was created by merging the TaoMate-H3 3-step LoRA directly into the official pruned MiniMax H3 FL2VA INT8 ConvRot-256 weights.
Official TaoMate-H3 project:
https://github.com/TaoLiveAIGC/TaoMate-H3
The original TaoMate adapter uses:
- LoRA rank: 128
- LoRA alpha: 128
- Native alpha/rank scale: 1.0
- Merge strength used for this release: 0.9
The LoRA contains:
- 416 LoRA tensors
- 208 target modules
Target distribution:
- 200 main transformer targets
- 8 Token Refiner targets
The merge includes the Token Refiner LoRA targets as well as the main transformer targets.
Merge Method
The official pruned INT8 ConvRot-256 MiniMax H3 FL2VA model was used as the starting point.
The TaoMate LoRA delta was calculated from the low-rank matrices and merged into the corresponding base weights.
For the main INT8 ConvRot layers, the existing quantized representation was reconstructed, the LoRA delta was applied in BF16 precision, and the resulting weights were requantized to INT8 ConvRot-256 using row-wise scaling.
The Token Refiner remained in BF16 and its eight TaoMate target modules were merged directly into the BF16 weights.
AdaLN parameters remained in FP16/high precision.
The final result is a single approximately 19.5 GB SafeTensors model with the TaoMate adaptation already integrated.
Why This Model Exists
The purpose of this release is to provide a practical single-file MiniMax H3 FL2VA model for ComfyUI that combines:
- MiniMax H3 audio-video generation
- TaoMate low-step adaptation
- INT8 ConvRot-256 quantization
- Integrated TaoMate LoRA
- No separate LoRA loading
- Low-step inference
- Audio and video generation in one model
This is particularly useful for ComfyUI users who want TaoMate-style low-step generation without loading the large TaoMate LoRA separately for every generation.
Important: This merged model contains the TaoMate LoRA weights, but it does not reproduce every component of the official TaoMate streaming runtime. The official TaoMate runtime contains additional streaming and KV-cache scheduling logic.
Recommended Settings
After testing multiple configurations, the recommended configuration for this merged model is:
| Parameter | Recommended |
|---|---|
| Steps | 8 |
| Sampler | Euler |
| Scheduler | Simple |
| CFG | 1.0 |
| Video Sigma Shift | 8 |
| Audio Sigma Shift | 3 |
| Denoise | 1.0 |
| Separate TaoMate LoRA | No |
Recommended Preset
Steps: 8
Sampler: Euler
Scheduler: Simple
CFG: 1.0
Video Sigma: 8
Audio Sigma: 3
Denoise: 1.0
The TaoMate LoRA is already merged into this model.
Do not load the original TaoMate LoRA again.
Why 8 Steps?
The original TaoMate adaptation is a 3-step distillation.
However, testing this merged ComfyUI model showed that 3 sampling steps were not optimal for visual quality.
Observed behavior:
3 steps
β visible instability in facial details, eyes, motion and fine structures
6 steps
β significantly improved, but some residual defects remained
8 steps
β noticeably more complete detail reconstruction and better structural stability
Therefore this release is specifically optimized for 8 sampling steps.
The underlying adaptation remains the TaoMate 3-step LoRA.
Sigma Shift Testing
Three test videos were generated to compare different Video Sigma Shift values.
The Audio Sigma Shift was kept at 3.
Video Sigma Shift 12 / Audio 3
Video Sigma Shift 6 / Audio 3
Video Sigma Shift 8 / Audio 3 β Recommended
The 8/3 configuration provided the best observed balance between:
- facial stability
- eye detail
- fine structures
- motion consistency
- overall visual quality
Therefore:
Video Sigma Shift = 8
Audio Sigma Shift = 3
Benchmark
Benchmark conditions:
- Video duration: 5 seconds
- Resolution: approximately 1 MP
- Sampling steps: 8
- ComfyUI generation
- Same local hardware/environment for both tests
This merged model
minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors
Generation time:
225 seconds
Reference configuration
minimax_h3_fl2va_pruned_int8_convrot.safetensors
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Generation time:
239 seconds
Difference
225 seconds vs 239 seconds
Difference: 14 seconds
Relative difference: approximately 5.9%
The merged model completed this particular benchmark approximately 5.9% faster than the tested INT8 + separate BF16 8-step reference configuration.
This is a local benchmark and should not be considered a universal performance guarantee. Performance can vary depending on GPU, VRAM, CUDA, PyTorch, ComfyUI version, resolution, workflow and other runtime conditions.
Benchmark Comparison
The comparison shows this merged model against:
minimax_h3_fl2va_pruned_int8_convrot.safetensors
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors
Example Videos
Sigma Shift comparisons
Performance comparison
Quantization
Main transformer:
Quantization: INT8
Format: ConvRot
Group size: 256
Scaling: Row-wise
The model preserves higher precision for components where maintaining numerical behavior is important:
AdaLN: FP16
Token Refiner: BF16
The Token Refiner is intentionally not represented as INT8 ConvRot in this release.
Model Structure
MiniMax H3 FL2VA
β
βββ Main Transformer
β βββ 50 transformer blocks
β βββ INT8 ConvRot-256
β βββ TaoMate merged into 200 target modules
β
βββ Token Refiner
β βββ 2 blocks
β βββ BF16
β βββ TaoMate merged into 8 target modules
β
βββ AdaLN
β βββ FP16
β
βββ Other H3 conditioning / patch components
Total TaoMate integration:
LoRA tensors: 416
Target modules: 208
Main targets: 200
Token Refiner: 8
Rank: 128
Alpha: 128
Merge strength: 0.9
Recommended ComfyUI Configuration
Model:
minimax_h3_fl2va_pruned_int8_convrot_taomate_8_step.safetensors
Steps:
8
Sampler:
Euler
Scheduler:
Simple
CFG:
1.0
Video Sigma Shift:
8
Audio Sigma Shift:
3
Denoise:
1.0
No additional TaoMate LoRA is required.
Base Model
Comfy-Org MiniMax H3:
https://huggingface.co/Comfy-Org/MiniMax-H3
Official MiniMax H3:
https://huggingface.co/MiniMaxAI/MiniMax-H3
TaoMate-H3:
https://github.com/TaoLiveAIGC/TaoMate-H3
Checksums
Original Comfy-Org pruned INT8 ConvRot base:
SHA256:
e889202c41dafb67b10d67b97f0d8541508036a6090af23425a5c2615d03c47a
This merged release:
SHA256:
302041433520d0ae7386fbe380cbd439d296b7ff9b1217537a37921e72c1e407
License
This model is a community derivative/merged release based on MiniMax H3 and TaoMate-H3 components.
Users must review and comply with the applicable licenses of the base model and the TaoMate-H3 project before redistribution or commercial use.
The Comfy-Org MiniMax H3 distribution identifies the MiniMax H3 Community License Agreement as its applicable license.
Disclaimer
This is an experimental community merge optimized and tested for ComfyUI.
The recommended 8-step / 8-3 configuration is based on empirical testing of this specific merged release.
Results may vary depending on hardware, software versions, workflow, resolution, generation length and other runtime parameters.
- Downloads last month
- 50