Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 10 days ago • 50
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 8 days ago • 114
CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments Paper • 2609.38087 • Published 8 days ago • 24
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents Paper • 2609.33772 • Published 10 days ago • 35
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning Paper • 2609.31199 • Published 12 days ago • 14
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 14 days ago • 24
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 14 days ago • 64
HuRo: Robotizing Human Videos for Scalable VLA Pretraining Paper • 2609.10706 • Published 19 days ago • 29
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 20 days ago • 93
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Paper • 2609.25165 • Published 16 days ago • 77
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 16 days ago • 157
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 19 days ago • 151
CADWorld: Computer-Use Benchmark for Long-Horizon Computer-Aided Design Paper • 2609.16251 • Published 23 days ago • 14
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 20 days ago • 224
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 20 days ago • 57
UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation Paper • 2609.12397 • Published 20 days ago • 46