VDN-H3
8-step hybrid-attention video + soundtrack, MiniMax-H3
None defined yet.
8-step hybrid-attention video + soundtrack, MiniMax-H3
Segment indoor 3D scans into object instances
Ask who said what in a multi-speaker recording
Editable HTML/CSS graphic design from a text brief
Drive a world model with WASD and camera keys
Low-latency streamable 4x video super-resolution
1930s hand-painted animation style, video+audio
Few-step MiniMax-H3 video with synchronized audio
Reference-conditioned video with sound, from MiniMax-H3
Expand ideas into structured MiniMax-H3 A/V prompts
Multilingual TTS and voice cloning with ICE-012-Audio
Re-rank passages with cross-encoders from arXiv 2603.03010
Restore damaged music in MiniMax Music 3 DAV latent space
Lightweight autoregressive TTS with voice cloning
Multimodal speech & image-grounded translation, 25 languages
Zero-shot voice-cloning TTS at 48 kHz from dots.tts
Turn a course brief into slides or an interactive widget
Egocentric video + instruction to robot trajectory
Generate and outpaint seamless 360° panoramas with Krea 2
Same prompt, many seeds - diversity LoRA for Krea 2 Turbo
Verbatim & intended ASR with word timestamps
Recurrent feed-forward 3D Gaussian Splatting
Thai document/handwriting OCR with OpenThai 2.0 27B VLM
Hebrew text to IPA phonemes for TTS