Avoid ZeroGPU startup OOM with precompiled Wan AOTI blocks

#3
by evalstate HF Staff - opened
MCP Tools org

Replace startup-time compilation of the full unquantized Wan pipeline with precompiled ZeroGPU AOTI transformer blocks. The application still loads and fuses the same two Lightx2v adapters at scales 3.0 and 1.0, then applies the existing int8/fp8 quantization before loading the published fp8da blocks. This prevents ZeroGPU from moving the oversized unquantized compilation closure to the worker while preserving the model, adapters, and generation interface.

Automated repair proposed by the scheduled Space monitor for revision 443882174ab7281cfdd6f93d1e7a78b182557e2d.
Findings: Actual Hub runtime evidence shows startup fails when optimize_pipeline_ invokes its decorated compile_transformer callback. ZeroGPU attempts to move the callback’s unquantized Wan pipeline to the A10G…; The candidate preserves both Lightx2v adapters and their original fusion scales, quantizes the text encoder and transformers before GPU dispatch, and loads the published Wan fp8da AOTI blocks instead …; The AOTI approach is supported by the canonical upstream Space and another LoRA-preserving Wan deployment currently running on ZeroGPU A10G. The current Spaces releases expose aoti_blocks_load, the pi…; The unauthenticated Hub run-log endpoint returned 401, so the full stream could not be read. Diagnosis used the actual traceback in runtime.raw.errorMessage; full end-to-end generation cannot be exerc…; All Python files in the complete verified source tree compile in memory. Static checks confirm the failing optimize_pipeline_ startup call is removed, the original adapter operations remain, quantizat…

evalstate changed pull request status to merged

Sign up or log in to comment