- MiniMax H3 X2 Detail VAE
- A Failed HQ-VAE Experiment That Became a 2-in-1 Release
- What a long chain of experiments taught us about latent bottlenecks, spatial detail, decoder capacity, and a misleadingly successful side-channel
- What Is Released
- Abstract
- 1. Starting Point
- 2. The First Trap: The Visible Grid
- 3. Decoder-Side Experiments
- 4. The Missing-Detail Target Problem
- 5. Can Late Decoder Features Predict the Missing Structure?
- 6. Searching for a Latent Detail Direction
- 7. Moving Backward Through the Encoder
- 8. Feature-Space Information Survival
- 9. The 256-to-512 Experimental Trap
- 10. Layerwise Encoder Predictability
- 11. Compressing Down1
- 12. The First Real Visual Success: B16
- 13. More Capacity Was Not the Answer
- 14. B32 Became the Detail Sweet Spot
- 15. Detail Was Not Just High Frequency
- 16. Spatial Gating Did Not Help
- 17. Localizing the Best Internal Tap
- 18. Full-Resolution Reconstruction
- 19. What We Had Actually Built
- 20. The Fatal Problem
- 21. Why This Fails for Text-to-Video
- 22. One Checkpoint File Does Not Fix the Data Flow
- 23. Distilling B32 from Information Available at Decode Time
- 24. The Generative Detail Prior
- 25. Why a Larger Decoder Cannot Recover Information It Never Received
- 26. What the Teacher Really Is
- 27. Runtime Integration Also Exposed a Practical Trap
- 28. The Most Dangerous Intermediate Result
- 29. Lessons From the Failed Branches
- 30. Final Outcome
- 31. What Would Need to Change for a True Higher-Information H3 Pipeline?
- 32. Future Experiment: Scale-Paired Latents
- A Failed HQ-VAE Experiment That Became a 2-in-1 Release
- Conclusion
MiniMax H3 X2 Detail VAE
Repository: speach1sdef178/MiniMax-H3-X2-Detail-VAE
A Failed HQ-VAE Experiment That Became a 2-in-1 Release
What a long chain of experiments taught us about latent bottlenecks, spatial detail, decoder capacity, and a misleadingly successful side-channel
Status: Working model + tested custom node + workflow + research report
Starting point: my own already-working MiniMax H3 2X VAE
Scope: this report intentionally begins with that completed 2X VAE. I created the base 2X VAE used for this project as well, but how that VAE was built is a separate story and is not covered here.
What Is Released
This repository is the final practical outcome of the experiment. It contains a 2-in-1 MiniMax H3 X2 VAE, an optional ComfyUI detail-enhancement node, an example workflow, and the full research article below.
The same model has two distinct uses:
Mode 1 β 2X VAE Decode
Load MiniMax-H3-X2-Detail-v1.safetensors with the normal ComfyUI Load VAE node and use it as the video VAE for MiniMax H3. For the actual video decode, replace the base ComfyUI VAE Decode node with MiniMax H3 VAE Decode (fast) from TripleHeadedMonkey/ComfyUI-MiniMaxH3_LatentUpscaler. This node is required for the released 2X workflow.
The validated starting settings shown in the included workflow are tiling = true, tile_size = 256, tile_overlap = 64, output_device = cpu, and temporal_tiling = false.
Generated H3 video latent
β
MiniMax H3 VAE Decode (fast) β MiniMax-H3-X2-Detail-v1
β
packed 12-channel X2 decode
β
PixelShuffle Γ2
β
2X RGB / video frames
Mode 2 β Reference Detail Enhancement
For image-reference workflows, place the included MiniMax H3 VAE 2X (Detailed Upscale) node between the source/reference image and MiniMax H3 Reference to Video. Select the same MiniMax-H3-X2-Detail-v1.safetensors model inside the node.
Reference image
β
MiniMax H3 VAE 2X (Detailed Upscale)
β
X2 detail-enhanced reference
β
MiniMax H3 Reference to Video
This second mode activates the additional early-encoder B32 spatial-detail path described in the article. It is not a latent-only enhancement path: it requires an existing RGB reference because the B32 representation is extracted from down.1.block.1 before the standard H3 latent bottleneck.
The two modes should therefore not be confused:
| Mode | Input | Custom node | Purpose |
|---|---|---|---|
| 2X VAE | H3 latent | No B32 node | 2X decode through MiniMax H3 VAE Decode (fast) |
| Detail Enhancer | RGB reference image | Yes | Reconstruct additional spatial structure before Reference-to-Video |
Included Files
The release package is prepared for the following files:
MiniMax-H3-X2-Detail-v1.safetensors # add separately: model is too large for this preparation archive
custom_node/
ComfyUI-H3-X2-Detailed.zip
workflow/
H3_VAE_Detailed_and_2X_VAE_2in1.json
README.md # this article / model card
The example workflow demonstrates both roles of the same model: it is loaded as the X2 video VAE and decoded through MiniMax H3 VAE Decode (fast), while the same checkpoint is also selected inside the Detailed Upscale node for optional reference enhancement.
Requirements and Installation
The 2X decode workflow requires TripleHeadedMonkey/ComfyUI-MiniMaxH3_LatentUpscaler, specifically its MiniMax H3 VAE Decode (fast) node. The included B32 reference-enhancement mode additionally uses ComfyUI-H3-X2-Detailed.zip from this release.
- Install
ComfyUI-MiniMaxH3_LatentUpscalerunderComfyUI/custom_nodes/and restart ComfyUI. - Put
MiniMax-H3-X2-Detail-v1.safetensorsinComfyUI/models/vae/. - Extract
ComfyUI-H3-X2-Detailed.zipintoComfyUI/custom_nodes/and restart ComfyUI. - Load the included workflow.
- For Mode 1, select
MiniMax-H3-X2-Detail-v1.safetensorsin Load VAE and connect that VAE toMiniMax H3 VAE Decode (fast)instead of the base VAE Decode node. - For Mode 2, additionally connect the source image through
MiniMax H3 VAE 2X (Detailed Upscale)beforeMiniMax H3 Reference to Video.detail_strength = 1.0is the validated baseline. Final video decode still usesMiniMax H3 VAE Decode (fast)with the same X2 VAE.
Included Workflow
The workflow demonstrates both uses of the same model in one graph: the green Detailed Upscale node enhances the RGB reference before Reference-to-Video, while the generated H3 latent is decoded at the end through MiniMax H3 VAE Decode (fast) using MiniMax-H3-X2-Detail-v1.safetensors.
Video Comparisons
Both comparisons show the same model in its two released operating modes:
Left β 2X VAE Mode
Right β 2X VAE + Detailed Mode
The generation settings are matched. In Detailed Mode, the reference is first processed through the B32 detail-enhancement path; final video decoding in both cases uses the same X2 VAE.
Comparison 01
Comparison 02
Tested reference-enhancement path: the included custom node has been tested successfully in the supplied workflow. Its architectural limitation remains important: the B32 path uses early encoder features from an existing RGB reference and should not be described as adding B32 detail to arbitrary generated H3 latents.
License
This release is derived from MiniMax H3 and is distributed subject to the MiniMax H3 Community License Agreement. Hugging Face metadata therefore uses license: other. Please read the LICENSE file and the accompanying NOTICE before downloading, redistributing, or using the model.
This repository contains modified VAE weights and additional custom components created for this project. The original MiniMax H3 project and its licensors are not affiliated with or responsible for these modifications.
Abstract
The goal sounded simple: take an already functional MiniMax H3 2X Video VAE and make the decoded result contain more genuine spatial information, not merely more pixels.
The starting 2X VAE could already decode at twice the spatial resolution. The problem was that the larger output did not contain a proportional increase in useful detail. Fine structures remained weak, and in many cases the result looked like a larger representation of essentially the same information.
We investigated token-to-patch assembly, PixelShuffle geometry, periodic grid artifacts, decoder-side spatial processing, missing-detail targets, latent detail directions, pre-bottleneck encoder representations, compact side-channels, decoder capacity, full-resolution reconstruction, deterministic distillation, internal decoder features, and generative detail priors.
One branch eventually produced a convincing improvement. A compact
32-channel representation extracted from down.1.block.1 of the frozen
H3 encoder restored visible local structures that the normal VAE path
did not reconstruct. On the full-resolution holdout test it improved
RMSE, HP3, HP5 and gradient metrics on all 10 tested frames.
For a moment, it looked like the project had succeeded.
It had not.
The successful system relied on information extracted from the original RGB image before the normal H3 latent bottleneck. That information is unavailable when MiniMax H3 creates a new latent during text-to-video generation. The result was therefore not a drop-in improved VAE decoder. It was an encoder-to-decoder side-channel.
This report documents how we arrived there, why several apparently promising intermediate results were misleading, and why the final negative result changed our understanding of the problem.
1. Starting Point
We began with an already working 2X MiniMax H3 VAE. The mature starting checkpoint used a 12-channel packed decoder output followed by external 2x PixelShuffle:
H3 latent
β
ViT3D decoder
β
proj_out
β
12 packed channels
β
PixelShuffle Γ2
β
2X RGB
The system already produced a valid image at twice the spatial resolution.
The research question was therefore not:
Can we make the VAE output a larger image?
It was:
Can we make the larger reconstruction contain additional meaningful spatial detail?
That distinction shaped the rest of the project.
2. The First Trap: The Visible Grid
The 2X output contained a conspicuous repeating spatial pattern. It was tempting to assume that the grid itself was the main reason fine detail was poor.
We reproduced the final projection and patch assembly exactly. The
replay matched bit-for-bit (mean L1 = 0, max L1 = 0), which allowed
us to inspect the geometry without introducing diagnostic reconstruction
errors.
The decoder structure turned out to be important:
- all 36 transformer blocks operate on the latent token grid;
proj_outexpands each token into a packed patch in one step;- there is no progressive high-resolution decoder hierarchy after the transformer;
- external PixelShuffle x2 converts the packed output to RGB.
Boundary-energy probes showed strong boundaries at approximately 16x16 in packed space and 32x32 in final X2 output space.
The grid was therefore closely related to the native token-to-patch reconstruction geometry.
But subsequent experiments showed that grid suppression and detail reconstruction were not the same problem. We eventually stopped treating grid removal as the primary objective.
3. Decoder-Side Experiments
The next hypothesis was that the normal latent contained enough information but the X2 decoder failed to use it.
We tried several decoder-side approaches:
- spatial detail blocks;
- token-spatial residual blocks;
- processing before
proj_out; - continuous spatial decoding across token boundaries;
- stronger progressive RGB decoders;
- residual reconstruction heads.
The motivation was reasonable: allow information to move spatially across patch boundaries instead of letting each latent token independently expand into its final patch.
In practice, the results fell into four categories:
- almost no visible change;
- stronger contours without new structure;
- oversharpening or severe artifacts;
- fragile training behaviour inside the frozen H3 decoder.
Increasing decoder sophistication did not demonstrate that the missing information was actually present in the decoder input.
4. The Missing-Detail Target Problem
A recurring problem was defining "detail."
Targets such as:
HQ - reconstruction
or high-pass versions of that residual were mathematically convenient, but a network can reduce those losses by learning edges, contrast changes or contour enhancement.
Several branches improved numerical objectives while producing:
- edge halos;
- engraving-like structure;
- oversharpening;
- nearly empty residuals;
- changes that were measurable but not useful.
This became a central rule for the project:
A loss named "detail loss" does not prove that the model is learning detail.
Visual inspection remained the primary gate.
5. Can Late Decoder Features Predict the Missing Structure?
Before building another large decoder, we tested whether the representation immediately before the final projection could predict the signed missing HQ microstructure.
The result was weak.
Late decoder features were useful for semantic and edge-related information, but they did not provide convincing evidence that the exact missing fine structure was recoverable.
That changed the question:
Instead of asking how to decode the information better, we needed to ask whether the information survived into the standard latent path at all.
6. Searching for a Latent Detail Direction
We also tested whether detail behaved like a stable direction in latent space.
Related degraded versions sometimes produced surprisingly low-dimensional latent changes. This suggested a possible axis:
less detail β normal β more detail
However, amplifying that direction mostly strengthened edges and produced engraving-like structure.
A stable latent direction was not necessarily a true detail direction. It could encode blur severity, contrast, frequency balance or another degradation property.
This branch was closed.
7. Moving Backward Through the Encoder
The encoder investigation changed the project.
The relevant hierarchy was approximately:
RGB
β
down0
β
down1
β
down2
β
down3
β
down4
β
down5
β
norm_out
β
SiLU
β
conv_out (48 channels)
β
first 24 channels
β
normalized H3 latent
An important architectural correction was that conv_out produces 48
channels. The deterministic encode path then selects the first 24
channels before the final per-channel normalization.
The final normalization is an invertible affine transform and should not itself be interpreted as an information-destroying bottleneck.
The meaningful loss occurred earlier.
8. Feature-Space Information Survival
Real-structure paired tests showed that earlier encoder representations distinguished fine structural differences much more strongly than the final 24-channel representation.
A TRAIN-only PCA side-channel gave approximately:
Compact representation HOLDOUT feature retention
16 channels ~92% 32 channels ~96% 64 channels ~98%
This was encouraging, but it exposed another trap:
Feature-space distinguishability is not the same as RGB decodability.
A representation can separate two inputs mathematically without allowing the missing image structure to be reconstructed accurately.
Matched reconstruction experiments confirmed this. A compact
pre-bottleneck representation did not automatically beat mean24 when
asked to reconstruct signed microstructure.
9. The 256-to-512 Experimental Trap
During this work we discovered that some earlier diagnostics used a 512x512 HQ target while the encoder had received a 256x256 downsampled image.
Those experiments were therefore measuring two things at once:
- information surviving through the encoder;
- super-resolution predictability from 256 to 512.
This distinction is critical. A model can predict plausible high-resolution structure that was never explicitly present in its low-resolution input.
We kept this limitation in mind when interpreting later layerwise results.
10. Layerwise Encoder Predictability
Despite that caveat, the layerwise trend was extremely clear.
Representative HOLDOUT medians were:
Representation RMSE Cosine
down0 0.0383 0.9458 down1 0.0404 0.9346 down2 0.0532 0.8802 down3 0.0626 0.8010 down4 0.0693 0.7708 norm 0.0695 0.7686 mean24 0.0705 0.7632
The useful signal degraded progressively as spatial resolution collapsed.
down1 was especially attractive: it retained high spatial resolution
while already containing learned H3 visual features.
The extraction point was localized to:
down.1.block.1
11. Compressing Down1
The down1 representation has 256 channels. We tested learned 1x1
channel bottlenecks:
256 β B8
256 β B16
256 β B32
256 β B64
Even B8 retained surprisingly useful information.
B16 was initially attractive because it represented a 16x reduction in channel count with only a modest loss of predictive quality.
This suggested that the useful spatial signal in down1 was highly
redundant in channel space.
12. The First Real Visual Success: B16
We froze the Stock H3 encoder and the E34 X2 carrier.
Only two components were trained:
down1 256 β 16
+
small spatial detail decoder
The branch predicted a bounded RGB residual that was added to E34.
On HOLDOUT, the approximate ratios relative to E34 were:
Metric Ratio
RGB RMSE 0.982 HP3 0.915 HP5 0.828 Gradient 0.887
19/20 crops improved in RGB RMSE, HP3 and HP5.
More importantly, the visual differences were meaningful:
- small flower elements became better separated;
- thin grasses and stems became more visible;
- textured surfaces recovered localized structure;
- some fur and fine boundaries became better defined.
This was the first branch that looked like additional spatial information rather than generic sharpening.
13. More Capacity Was Not the Answer
We tested several obvious extensions.
Down1 + Down2 fusion
Adding a second encoder stage did not meaningfully improve the result.
Progressive Spatial Decoder
A substantially larger decoder failed to beat the simpler B16 branch. On HP5 it was worse on all 20 HOLDOUT samples.
This was strong evidence that decoder capacity was not the primary bottleneck.
14. B32 Became the Detail Sweet Spot
Using the same decoder architecture, we compared B16, B32 and B64.
Relative to B16, B32 slightly worsened raw RGB RMSE but improved fine-detail measures:
Metric B32 / B16
RGB RMSE 1.0047 HP3 0.9908 HP5 0.9808 Gradient 0.9936
B32 won HP5 on 19/20 HOLDOUT crops.
B64 did not continue the improvement.
B32 therefore became the practical sweet spot for the detail branch.
This also demonstrated that pixel fidelity and fine structural reconstruction can conflict.
15. Detail Was Not Just High Frequency
A frequency-selective B32 variant attempted to isolate only high-frequency information.
It performed worse than the unrestricted B32 branch and won HP5 on only 3/20 HOLDOUT cases.
The implication was important:
Useful small structures require medium-frequency spatial support as well as high-frequency components.
A blade of grass or a flower petal is not simply "high-frequency energy." It has shape.
16. Spatial Gating Did Not Help
We trained a spatial gate:
Final = E34 + gate(x,y) Γ B32 residual
The learned gate remained almost fully open (mean β 0.999).
The output was effectively the same as the ungated B32 result.
The B32 residual did not generally need aggressive post-suppression.
17. Localizing the Best Internal Tap
We compared compatible internal representations around the successful
down1 point.
down.1.block.0 was clearly worse on fine-detail metrics than
down.1.block.1.
The useful path became tightly localized:
down.1.block.1
β
256 β B32
β
small spatial decoder
β
bounded RGB residual
β
E34 + residual
At this stage, further micro-optimization around the same local architecture no longer looked productive.
18. Full-Resolution Reconstruction
The next question was whether B32 only worked on small crops.
Full-resolution extraction introduced a practical complication: H3's adaptive encoder internally tiles large images, causing a normal forward hook to fire many times.
We therefore explicitly extracted down1 using 256x256 X1 tiles with
overlap 64 and stitched the feature maps using overlap averaging.
The normal H3 latent was produced separately through the canonical encode path.
Across 10 full HOLDOUT frames:
Metric Median ratio vs E34 Wins
RGB RMSE 0.9889 10/10 HP3 0.9279 10/10 HP5 0.8464 10/10 Gradient 0.8864 10/10
Median residual RMS was only about 0.00724.
Visually, the branch restored thin grass structures, localized surface detail, fur boundaries and small structural separations without obvious repeated texture stamps or catastrophic seams in the inspected frames.
At this point the project looked successful.
19. What We Had Actually Built
The successful system was:
Stock H3 Encoder
RGB ββββββββββββββββββββββββββββ¬ββββββββββββββββββββ
β
ββββββββββββ΄βββββββββββ
β β
βΌ βΌ
down.1.block.1 normal H3
256 channels 24ch latent
β β
βΌ βΌ
B32 E34 X2
β decoder
βΌ β
spatial decoder RGB
β β
βΌ β
RGB residual ββββββββββββββββ
β
βΌ
Final X2 RGB
The B32 branch was effectively a compact encoder-to-decoder bypass.
That fact became fatal to the original objective.
20. The Fatal Problem
A conventional VAE decoder receives:
latent β RGB
Our successful system received:
latent
+
early encoder features from the original RGB
β RGB
B32 worked precisely because it bypassed the normal H3 latent bottleneck.
It did not prove that the standard 24-channel H3 latent contained enough information for the improved reconstruction.
It proved that useful information existed before the bottleneck and could be preserved through a separate side-channel.
21. Why This Fails for Text-to-Video
For reconstruction of an existing image, we can compute:
RGB β down1 β B32
For T2V there is no original RGB image.
The generation path is:
noise
β
MiniMax H3 diffusion
β
generated 24ch H3 latent
β
VAE decoder
β
video
At decode time, the generated latent exists.
The original down.1.block.1 features do not.
The Teacher's most important input is therefore unavailable.
This is why the successful reconstruction system cannot act as a true drop-in higher-detail T2V VAE.
22. One Checkpoint File Does Not Fix the Data Flow
We packaged E34 and the B32 branch into one .safetensors checkpoint.
That does not change the architecture.
The file can contain:
E34 weights
+
B32 projection
+
B32 decoder
but the B32 branch still requires down.1.block.1 features at runtime.
A single checkpoint file does not imply a conventional latent-to-RGB model.
The runtime information path matters more than packaging.
23. Distilling B32 from Information Available at Decode Time
We then tried to eliminate the privileged RGB side-channel.
E34 RGB β B32 residual
A deterministic student was trained to predict the Teacher's B32 residual from E34 RGB alone.
On HOLDOUT it reached approximately:
RMSE β 0.00862
cosine β 0.653
pred/target RMS β 0.629
Some samples were predicted very well, but the student did not reproduce the full Teacher signal.
Internal E34 decoder features
We added internal decoder information, including the 2048-dimensional
pre-proj_out representation.
It did not beat the RGB-only student.
This suggested that the missing Teacher information was not simply hidden in a convenient late-decoder representation.
24. The Generative Detail Prior
If exact information was unavailable, perhaps a model could generate plausible missing detail.
The first generative prior produced obvious dirt/noise.
A second version separated deterministic structure from a stochastic teacher-minus-student residual and used learned support/amplitude conditioning.
It was cleaner, but still worse than the deterministic student.
Representative HOLDOUT results:
Variant RMSE Cosine
deterministic RGB student 0.00862 0.653 support-conditioned generative version 0.00901 0.571
Increasing stochastic strength progressively degraded the result.
On a nearly perfect control sample, the true Teacher residual RMS was
only about 0.00075, while one generative configuration produced
roughly 0.0097.
The model was converting uncertainty into texture.
Visually, the effect looked like sand or dirt.
The generative branch was closed.
25. Why a Larger Decoder Cannot Recover Information It Never Received
This became the central conclusion.
If two high-resolution structures collapse to sufficiently similar standard latents, a deterministic decoder cannot know which structure was originally present.
A larger decoder can produce a more plausible guess:
latent β plausible hair
latent β plausible grass
latent β plausible fabric
but that is generation, not faithful reconstruction.
Our experiments repeatedly exposed the difference.
The B32 Teacher worked because it did not have to guess: it received information from before the bottleneck.
26. What the Teacher Really Is
The Teacher should not be described as a normal improved VAE.
A better description is:
A compact high-spatial-resolution encoder-to-decoder side-channel attached to an X2 H3 reconstruction path.
B32 transforms:
down1:
256 Γ H/2 Γ W/2
into:
B32:
32 Γ H/2 Γ W/2
This is an 8x channel reduction while retaining useful spatial information.
The RGB residual decoder is only one possible consumer of that representation.
The experiment therefore failed as a drop-in Video VAE but produced an interesting compact H3-native spatial representation.
Possible future research directions include reference representation, spatial guidance, correspondence, feature propagation and temporal side information. These applications were not demonstrated in this project and should be treated as future work rather than established results.
27. Runtime Integration Also Exposed a Practical Trap
Full-resolution B32 extraction uses tiled access to down.1.block.1.
Our first ComfyUI runtime node repeatedly entered the normal VAE model-management path while processing tiles. With DynamicVRAM enabled, this caused large amounts of repeated:
Model MiniMaxH3VideoVAE prepared for dynamic VRAM loading...
logging.
This was an implementation problem in the initial runtime node rather than
a failure of the B32 checkpoint itself. It was fixed in the released node
by preparing the H3 VAE once for the extraction pass and processing the
256x256 tiles through the already resident underlying encoder instead of
re-entering the public wrapper.encode() path for every tile.
The corrected node was then tested successfully in the supplied workflow. The repeated DynamicVRAM preparation behavior was no longer present, while the Teacher/B32 reconstruction path remained unchanged. This runtime issue is therefore documented here as one of the implementation pitfalls we met, not as a known problem of the released node.
28. The Most Dangerous Intermediate Result
The B32 full-resolution result was convincing:
- the images looked better;
- fine-detail metrics improved;
- the effect generalized beyond crops;
- the residual remained small;
- obvious repeated texture artifacts were absent.
It was very easy to conclude:
We built the HQ VAE.
But the successful experiment had silently changed the task.
Original goal:
H3 latent β better reconstruction
Successful experiment:
H3 latent
+
privileged pre-bottleneck information
β better reconstruction
Those are fundamentally different problems.
This may be the most useful general lesson from the project:
When reconstruction suddenly improves, verify that the model has not gained access to information that will be unavailable in the intended deployment path.
29. Lessons From the Failed Branches
**Metric improvement is not sufficient.**
Several branches improved numerical objectives while producing no
meaningful visible detail.
**High frequency is not synonymous with detail.**
The frequency-isolated branch was worse than unrestricted B32.
**More parameters do not create information.**
A stronger progressive decoder failed to beat the small B32 decoder.
**Feature-space retention does not guarantee visual decodability.**
PCA retention looked excellent but did not automatically translate into
RGB microstructure.
**Super-resolution predictability and information preservation are
different questions.**
A 512 target with a 256 encoder input mixes the two.
**A generative prior changes the problem.**
Once missing information is invented, the task is no longer faithful
reconstruction.
**A checkpoint file does not define the architecture.**
The actual runtime data flow does.
**Visual evaluation remained essential.**
The objective was real visible structure, not merely a lower proxy loss.
30. Final Outcome
The project progressed through:
working X2 VAE
β
patch/grid localization
β
decoder-side processing
β
detail-target experiments
β
decoder-token predictability
β
latent detail directions
β
encoder information survival
β
down1 localization
β
B16
β
B32
β
full-resolution success
β
deployment analysis
β
side-channel realization
β
latent-only distillation
β
internal-feature student
β
generative prior
β
negative result
The final conclusion is deliberately narrow:
We did not find a way to reproduce the B32 Teacher's real reconstruction improvement from the standard generated MiniMax H3 latent alone.
The best reconstruction result required information that bypassed the standard latent bottleneck.
Therefore it could not become the drop-in T2V/I2V Video VAE we originally intended to build.
31. What Would Need to Change for a True Higher-Information H3 Pipeline?
The Teacher suggests that a genuinely higher-information architecture would need to preserve more spatial information before the standard bottleneck.
Conceptually:
NEW encoder
β
βββββββββββββββββββββββββββ
β generative latent β
β + β
β compact spatial detail β
βββββββββββββββββββββββββββ
β
NEW generative model
β
NEW decoder
But MiniMax H3's diffusion model is trained to generate its existing latent representation, not an additional B32-like detail representation.
Changing this would therefore mean modifying more than the VAE decoder. It would affect the latent representation and the generative model itself.
That is a substantially different project.
32. Future Experiment: Scale-Paired Latents
One remaining idea is to compare the same image encoded at different spatial scales using the same H3 VAE.
For example:
512 image β H3 encoder β Z512
256 version β H3 encoder β Z256
After careful geometry matching or decode/re-encode canonicalization, the scale-dependent latent component could be compared against the real structures recovered by the B32 Teacher.
The key question would be:
Does the standard H3 latent contain a stable scale-dependent component that correlates with the real spatial detail preserved by B32?
If yes, there may still be a latent-only route worth investigating.
If not, it would further support the conclusion that the useful information disappears before the normal generated representation.
This experiment remains future work.
Conclusion
We tried to turn an already functional MiniMax H3 2X VAE into a genuinely higher-detail Video VAE.
We failed at that engineering objective.
The failure was nevertheless informative.
The visible grid was not the whole problem. Cross-patch processing did not solve it. Larger decoders did not solve it. High-frequency supervision did not solve it. Latent extrapolation did not solve it. Late decoder features did not solve it. A generative prior mostly turned uncertainty into artificial texture.
The first convincing improvement appeared only when we moved backward
through the encoder and preserved information from down.1.block.1
through a compact B32 side-channel.
That system worked because it bypassed the normal H3 latent bottleneck.
And that is exactly why it could not serve as a normal latent-only Video VAE for newly generated T2V frames.
The final result is therefore not a new drop-in HQ VAE.
It is evidence that a compact early-encoder representation can preserve spatial information that the standard H3 latent path does not retain well enough for our reconstruction objective.
Sometimes the most useful result of a failed decoder project is discovering that the decoder was not where the important information was being lost.
Release Notes
This repository intentionally publishes the practical artifacts together with the negative-result article. The original objective β a drop-in latent-only HQ VAE that reconstructs the B32 Teacher detail from an ordinary generated H3 latent β was not achieved.
However, the work produced two usable outcomes:
- the underlying X2 VAE remains usable as a normal MiniMax H3 2X VAE;
- the B32 branch can be used as an optional reference-image detail enhancer through the included custom node.
The model is therefore released as an experimental 2-in-1 tool, not as a claim that the latent bottleneck problem was solved.
Base X2 VAE provenance
The starting X2 VAE used throughout this research was also created by me. This article deliberately begins after that model was already working. The experiments and development process that produced the original X2 VAE are a separate story and are outside the scope of this report.
Recommended repository contents
MiniMax-H3-X2-Detail-v1.safetensorsβ model weights;custom_node/ComfyUI-H3-X2-Detailed.zipβ tested B32 reference-enhancement node;workflow/H3_VAE_Detailed_and_2X_VAE_2in1.jsonβ example ComfyUI workflow;assets/workflow_2in1.pngβ complete workflow screenshot;assets/minimax_h3_vae_decode_fast.pngβ required 2X decode node screenshot;assets/comparison_01.mp4andassets/comparison_02.mp4β two-mode visual comparisons (2X VAEvs2X VAE + Detailed);README.mdβ model card and full research article.
Intermediate checkpoints, training datasets and diagnostic scratch outputs are not required for normal use and are not part of the release package.
Disclaimer
This is an independent experimental research report. MiniMax H3 is referenced only to identify the architecture being investigated. The reported results describe the specific experimental setup used in this project and should not be generalized beyond it without independent validation.
- Downloads last month
- 15,466
Model tree for speach1sdef178/MiniMax-H3-X2-Detail-VAE
Base model
MiniMaxAI/MiniMax-H3
