DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Paper • 2608.31106 • Published 6 days ago • 98
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published Jul 28 • 66
Explicit Layer Modeling for Video Object Insertion and Layer Decomposition Paper • 2607.25802 • Published Jul 28 • 9
A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples Paper • 2607.29122 • Published Jul 31 • 5
Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations Paper • 2608.01628 • Published Aug 3 • 23
InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis Paper • 2608.02437 • Published Aug 3 • 72
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 183
Uncertainty-Aware World Model for Aerial Image-Goal Navigation Paper • 2608.05597 • Published Aug 6 • 12
Causal-JEPA: Learning World Models through Object-Level Latent Interventions Paper • 2602.11389 • Published Feb 11 • 13
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Paper • 2606.26907 • Published Jun 25 • 55
view article Article Diffusers welcomes FLUX-2 +6 YiYiXu, dg845, sayakpaul, OzzyGT, dn6, ariG23498, linoyts, multimodalart • Nov 25, 2025 • 190
Hala Technical Report: Building Arabic-Centric Instruction & Translation Models at Scale Paper • 2509.14008 • Published Sep 17, 2025 • 90
view article Article AtlasOCR: Building the First Open-Source Darija OCR Model with Vision Language Models imomayiz • Sep 16, 2025 • 19
view article Article Make your ZeroGPU Spaces go brrr with ahead-of-time compilation +2 cbensimon, sayakpaul, linoyts, multimodalart • Sep 2, 2025 • 82
view article Article Jupyter Agents: training LLMs to reason with notebooks +1 baptistecolle, hannayukhymenko, lvwerra • Sep 10, 2025 • 67